使用PyPDF2、PyMuPDF提取PDF文本出现乱码,原因是什么?
PDF解析乱码问题的解决方案
排查PDF类型:如果你的PDF是扫描生成的图片型PDF,常规文本提取工具(PyPDF2、PyMuPDF)无法直接提取文本,需要结合OCR识别:
安装依赖:pip install pytesseract pillow pymupdf,同时需安装Tesseract OCR引擎(Ubuntu:sudo apt install tesseract-ocr tesseract-ocr-chi-sim;Windows需下载安装包并配置环境变量)。
示例代码:import fitz import pytesseract from PIL import Image import io filename = "test2.pdf" with fitz.open(filename) as pdf: for page_num, page in enumerate(pdf): image_list = page.get_images(full=True) for img_index, img in enumerate(image_list): xref = img[0] base_image = pdf.extract_image(xref) image_bytes = base_image["image"] image = Image.open(io.BytesIO(image_bytes)) # 中文识别需指定lang='chi_sim',英文可省略 text = pytesseract.image_to_string(image, lang='chi_sim') print(f"第{page_num+1}页 图片{img_index+1}文本:\n{text}")换用更适配字体的工具:部分PDF因未嵌入字体导致提取乱码,可尝试使用pdfplumber,它对字体映射的处理更完善:
安装依赖:pip install pdfplumber
示例代码:import pdfplumber with pdfplumber.open("test2.pdf") as pdf: for page in pdf.pages: print(page.extract_text())调整PyMuPDF提取参数:尝试移除
sort=True,或改用按块提取的方式整理文本:filename = "test2.pdf" with fitz.open(filename) as f: for p in f: # 按块提取文本并排序 blocks = p.get_text("blocks") blocks.sort(key=lambda b: (b[1], b[0])) text = "\n".join([block[4] for block in blocks]) print("\n\n") print(text)
内容的提问来源于stack exchange,提问作者Hemil Parmar
相关产品推荐
相关产品推荐

