You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python批量转换docx到PDF脚本报错,如何修复?

批量将.docx文件转换为PDF的代码修复方案

问题描述

需要用Python批量将多个.docx文件转换为PDF,但运行代码时出现错误,错误信息如下:

Traceback (most recent call last):
  File "D:\Convert docx to pdf.py", line 26, in <module>
    document = Document()
NameError: name 'Document' is not defined

原代码:

import re
import os
from pathlib import Path
import sys
from docx2pdf import convert

# The location where the files are located
input_path = r'c:\Folder7\input'
# The location where we will write the PDF files
output_path = r'c:\Folder7\output'
# Creeaza structura de foldere daca nu exista
os.makedirs(output_path, exist_ok=True)

# Verifica existenta folder-ului
directory_path = Path(input_path)
if directory_path.exists() and directory_path.is_dir():
    print(directory_path, "exists")
else:
    print(directory_path, "is invalid")
    sys.exit(1)

for file_path in directory_path.glob("*"):
    # file_path is a Path object

    print("Procesez fisierul:", file_path)
    document = Document()
    # file_path.name is the name of the file as str without the Path
    document.add_heading(file_path.name, 0)

    file_content = file_path.read_text(encoding='UTF-8')
    document.add_paragraph(file_content)

    # build the new path where we store the files
    output_file_path = os.path.join(output_path, file_path.name + ".pdf")

    document.save(output_file_path)
    print("Am convertit urmatorul fisier:", file_path, "in: ", output_file_path)

修复方案

问题根源

  1. 未导入Document类,且该类属于python-docx库,用于创建/编辑docx文档,但这不是批量转换现有docx到PDF的正确方式。
  2. 原代码逻辑错误:没有读取现有docx文件,而是新建空白文档写入文件名和文本内容,完全偏离“转换现有docx”的需求。

修改后的代码

直接利用docx2pdf库的convert方法处理现有docx文件,同时过滤非docx格式的文件:

import os
from pathlib import Path
import sys
from docx2pdf import convert

# 输入文件夹路径
input_path = r'c:\Folder7\input'
# 输出PDF文件夹路径
output_path = r'c:\Folder7\output'

# 创建输出文件夹(不存在则创建)
os.makedirs(output_path, exist_ok=True)

# 验证输入文件夹是否有效
directory_path = Path(input_path)
if not (directory_path.exists() and directory_path.is_dir()):
    print(f"{directory_path} 无效")
    sys.exit(1)
print(f"{directory_path} 已存在")

# 遍历所有docx文件
for docx_file in directory_path.glob("*.docx"):
    print(f"正在处理文件: {docx_file}")
    # 构造输出PDF路径
    pdf_filename = docx_file.stem + ".pdf"
    pdf_file_path = os.path.join(output_path, pdf_filename)
    # 执行转换
    convert(docx_file, pdf_file_path)
    print(f"已转换文件: {docx_file} -> {pdf_file_path}")

说明

  • 使用glob("*.docx")确保只处理docx格式文件,避免无关文件干扰
  • docx_file.stem获取不带后缀的文件名,避免出现xxx.docx.pdf这类错误文件名
  • docx2pdf依赖Microsoft Word或LibreOffice(Windows下默认用Word,Linux/macOS需配置),确保系统已安装对应办公软件

内容的提问来源于stack exchange,提问作者Hellena Crainicu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 18:45:21