docx2pdf转换docx为PDF时提示文件损坏报错如何解决
问题现象
调用docx2pdf库批量转换指定目录下docx文件为PDF时,所有目标PDF文件可正常生成,但脚本运行结束前持续抛出COM组件错误,提示文件损坏。
原始问题代码
import os import sys from docx2pdf import convert convert("D:\Dev2Ease\PyC\Output Word","D:\Dev2Ease\PyC\Output Pdf") print ("All the files have been converted to pdf")
完整报错信息
Traceback (most recent call last): File "D:/Dev2Ease/PyC/word2pdf.py", line 5, in <module> convert("D:\Dev2Ease\PyC\Output Word","D:\Dev2Ease\PyC\Output Pdf") File "C:\Users\Shailesh\AppData\Roaming\Python\Python310\site-packages\docx2pdf\__init__.py", line 106, in convert return windows(paths, keep_active) File "C:\Users\Shailesh\AppData\Roaming\Python\Python310\site-packages\docx2pdf\__init__.py", line 25, in windows doc = word.Documents.Open(str(docx_filepath)) File "<COMObject <unknown>>", line 5, in Open pywintypes.com_error: (-2147352567, 'Exception occurred.', (0, 'Microsoft Word', 'The file appears to be corrupted.', 'wdmain11.chm', 25272, -2146822496), None)
问题产生原因
- 路径转义错误:Windows路径使用的反斜杠
\是Python字符串的默认转义符,直接书写未做特殊处理的路径字符串时,\加后续字母的组合会被解析为特殊转义字符,最终传给Word COM组件的文件路径存在异常,Word尝试打开无效路径对应的文件时就会抛出文件损坏错误。 - 临时文件被误识别:docx2pdf默认遍历输入目录下的所有文件,Word编辑文档时自动生成的以
~$开头的临时锁文件也会被识别为待转换的docx文件,这类临时文件本身不是完整的有效docx文档,打开时必然触发损坏报错。由于这类临时文件通常在目录遍历顺序的末尾,因此所有正常docx文件会先完成PDF转换,最后才抛出错误,就出现了“PDF全部生成成功但脚本报错”的现象。
修复方案
- 路径字符串统一加
r前缀声明为原始字符串,避免转义问题,也可以直接用正斜杠/书写Windows路径。 - 不要直接将目录路径传给convert函数,手动遍历输入目录,过滤掉临时文件、非docx后缀文件、子目录后,再对有效docx文件执行转换操作。
修复后的参考代码:
import os from docx2pdf import convert # 路径使用原始字符串,避免转义问题 input_dir = r"D:\Dev2Ease\PyC\Output Word" output_dir = r"D:\Dev2Ease\PyC\Output Pdf" # 提前创建输出目录,避免目录不存在报错 os.makedirs(output_dir, exist_ok=True) for file_name in os.listdir(input_dir): # 过滤Word临时文件、非docx后缀文件 if file_name.startswith("~$") or not file_name.lower().endswith(".docx"): continue file_full_path = os.path.join(input_dir, file_name) # 跳过子目录,仅处理文件 if os.path.isfile(file_full_path): convert(file_full_path, output_dir) print("All valid docx files have been converted to pdf")
内容的提问来源于stack exchange,提问作者shall
相关产品推荐
相关产品推荐

