You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyPDF2、PyMuPDF提取PDF文本出现乱码,原因是什么?

PDF解析乱码问题的解决方案
  • 排查PDF类型:如果你的PDF是扫描生成的图片型PDF,常规文本提取工具(PyPDF2、PyMuPDF)无法直接提取文本,需要结合OCR识别:
    安装依赖:pip install pytesseract pillow pymupdf,同时需安装Tesseract OCR引擎(Ubuntu:sudo apt install tesseract-ocr tesseract-ocr-chi-sim;Windows需下载安装包并配置环境变量)。
    示例代码:

    import fitz
    import pytesseract
    from PIL import Image
    import io
    
    filename = "test2.pdf"
    with fitz.open(filename) as pdf:
        for page_num, page in enumerate(pdf):
            image_list = page.get_images(full=True)
            for img_index, img in enumerate(image_list):
                xref = img[0]
                base_image = pdf.extract_image(xref)
                image_bytes = base_image["image"]
                image = Image.open(io.BytesIO(image_bytes))
                # 中文识别需指定lang='chi_sim',英文可省略
                text = pytesseract.image_to_string(image, lang='chi_sim')
                print(f"第{page_num+1}页 图片{img_index+1}文本:\n{text}")
    
  • 换用更适配字体的工具:部分PDF因未嵌入字体导致提取乱码,可尝试使用pdfplumber,它对字体映射的处理更完善:
    安装依赖:pip install pdfplumber
    示例代码:

    import pdfplumber
    
    with pdfplumber.open("test2.pdf") as pdf:
        for page in pdf.pages:
            print(page.extract_text())
    
  • 调整PyMuPDF提取参数:尝试移除sort=True,或改用按块提取的方式整理文本:

    filename = "test2.pdf"
    with fitz.open(filename) as f:
        for p in f:
            # 按块提取文本并排序
            blocks = p.get_text("blocks")
            blocks.sort(key=lambda b: (b[1], b[0]))
            text = "\n".join([block[4] for block in blocks])
            print("\n\n")
            print(text)
    

内容的提问来源于stack exchange,提问作者Hemil Parmar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 03:35:16