You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdf2docx转换PDF至Word时遇页面生成错误的解决需求

解决pdf2docx转换时「list index out of range」错误的方案

这个错误是pdf2docx解析复杂页面元素(如特殊布局的图片、非标准文本块)时触发的索引越界问题,转黑白无效说明问题不在颜色,而是页面结构解析bug。以下是几种可行方案:

方案1:修改pdf2docx参数,忽略图片并启用宽松解析模式

通过设置ignore_image=True跳过图片解析,同时用layout_recovery_mode="loose"降低布局解析的严格度,减少错误触发:

from pdf2docx import Converter

def pdf_to_word(pdf_path, word_output_path):
    cv = Converter(pdf_path, ignore_image=True)
    cv.convert(word_output_path, start=0, end=None, layout_recovery_mode="loose")
    cv.close()

if __name__ == "__main__":
    pdf_path = "你的输入文件路径.pdf"
    word_output_path = "你的输出文件路径.docx"
    pdf_to_word(pdf_path, word_output_path)

方案2:分页面处理,出错时自动 fallback 到纯文本提取

对每个页面单独尝试转换,失败时用pdfplumber提取纯文本,保证所有页面内容都能导出:

首先安装依赖:

pip install pdfplumber python-docx

代码实现:

from pdf2docx import Converter
from pdfplumber import Pdf
from docx import Document
import os

def process_single_page(pdf_path, page_num, doc):
    try:
        # 尝试用pdf2docx转换单页
        cv = Converter(pdf_path)
        temp_doc = "temp_page.docx"
        cv.convert(temp_doc, start=page_num, end=page_num+1)
        cv.close()
        # 合并临时文档内容到主文档
        temp_doc_obj = Document(temp_doc)
        for elem in temp_doc_obj.element.body:
            doc.element.body.append(elem)
        os.remove(temp_doc)
    except Exception as e:
        # 转换失败,提取纯文本
        with Pdf(pdf_path) as pdf:
            page = pdf.pages[page_num]
            text = page.extract_text()
            if text:
                doc.add_paragraph(text)
        print(f"第{page_num+1}页因错误{e},已用纯文本提取")

def pdf_to_word_with_fallback(pdf_path, word_output_path):
    doc = Document()
    with Pdf(pdf_path) as pdf:
        total_pages = len(pdf.pages)
    
    for page_num in range(total_pages):
        print(f"处理第{page_num+1}/{total_pages}页")
        process_single_page(pdf_path, page_num, doc)
    
    doc.save(word_output_path)

if __name__ == "__main__":
    pdf_path = "你的输入文件路径.pdf"
    word_output_path = "你的输出文件路径.docx"
    pdf_to_word_with_fallback(pdf_path, word_output_path)

方案3:纯文本提取生成Word(无格式,最稳定)

如果不需要保留原PDF的格式,直接用pdfplumber提取所有文本后写入Word,完全避开解析bug:

from pdfplumber import Pdf
from docx import Document

def pdf_text_to_word(pdf_path, word_output_path):
    doc = Document()
    with Pdf(pdf_path) as pdf:
        for page_num, page in enumerate(pdf.pages, 1):
            text = page.extract_text()
            if text:
                doc.add_paragraph(text)
                doc.add_page_break()
            print(f"已处理第{page_num}/{len(pdf.pages)}页")
    doc.save(word_output_path)

if __name__ == "__main__":
    pdf_path = "你的输入文件路径.pdf"
    word_output_path = "你的输出文件路径.docx"
    pdf_text_to_word(pdf_path, word_output_path)

内容的提问来源于stack exchange,提问作者Shady Mohamed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 23:19:57