Python批量转换docx到PDF脚本报错,如何修复?
批量将.docx文件转换为PDF的代码修复方案
问题描述
需要用Python批量将多个.docx文件转换为PDF,但运行代码时出现错误,错误信息如下:
Traceback (most recent call last): File "D:\Convert docx to pdf.py", line 26, in <module> document = Document() NameError: name 'Document' is not defined
原代码:
import re import os from pathlib import Path import sys from docx2pdf import convert # The location where the files are located input_path = r'c:\Folder7\input' # The location where we will write the PDF files output_path = r'c:\Folder7\output' # Creeaza structura de foldere daca nu exista os.makedirs(output_path, exist_ok=True) # Verifica existenta folder-ului directory_path = Path(input_path) if directory_path.exists() and directory_path.is_dir(): print(directory_path, "exists") else: print(directory_path, "is invalid") sys.exit(1) for file_path in directory_path.glob("*"): # file_path is a Path object print("Procesez fisierul:", file_path) document = Document() # file_path.name is the name of the file as str without the Path document.add_heading(file_path.name, 0) file_content = file_path.read_text(encoding='UTF-8') document.add_paragraph(file_content) # build the new path where we store the files output_file_path = os.path.join(output_path, file_path.name + ".pdf") document.save(output_file_path) print("Am convertit urmatorul fisier:", file_path, "in: ", output_file_path)
修复方案
问题根源
- 未导入
Document类,且该类属于python-docx库,用于创建/编辑docx文档,但这不是批量转换现有docx到PDF的正确方式。 - 原代码逻辑错误:没有读取现有docx文件,而是新建空白文档写入文件名和文本内容,完全偏离“转换现有docx”的需求。
修改后的代码
直接利用docx2pdf库的convert方法处理现有docx文件,同时过滤非docx格式的文件:
import os from pathlib import Path import sys from docx2pdf import convert # 输入文件夹路径 input_path = r'c:\Folder7\input' # 输出PDF文件夹路径 output_path = r'c:\Folder7\output' # 创建输出文件夹(不存在则创建) os.makedirs(output_path, exist_ok=True) # 验证输入文件夹是否有效 directory_path = Path(input_path) if not (directory_path.exists() and directory_path.is_dir()): print(f"{directory_path} 无效") sys.exit(1) print(f"{directory_path} 已存在") # 遍历所有docx文件 for docx_file in directory_path.glob("*.docx"): print(f"正在处理文件: {docx_file}") # 构造输出PDF路径 pdf_filename = docx_file.stem + ".pdf" pdf_file_path = os.path.join(output_path, pdf_filename) # 执行转换 convert(docx_file, pdf_file_path) print(f"已转换文件: {docx_file} -> {pdf_file_path}")
说明
- 使用
glob("*.docx")确保只处理docx格式文件,避免无关文件干扰 docx_file.stem获取不带后缀的文件名,避免出现xxx.docx.pdf这类错误文件名docx2pdf依赖Microsoft Word或LibreOffice(Windows下默认用Word,Linux/macOS需配置),确保系统已安装对应办公软件
内容的提问来源于stack exchange,提问作者Hellena Crainicu
相关产品推荐
相关产品推荐

