使用pdfminer与fitz处理Ghostscript生成PDF时遇Type3字体不支持错误
解决Ghostscript生成PDF的Type 3字体文本提取报错问题
核心问题背景
Ghostscript生成的PDF包含Type 3字体时,pdfminer.six(20221105版本)和fitz(PyMuPDF)默认会触发RuntimeError: pdf device does not support type 3 fonts,导致文本提取流程中断。以下是经过验证的解决和规避方案:
方案1:重新生成PDF时转换Type 3字体
直接通过Ghostscript参数,在生成PDF阶段将Type 3字体转换为工具支持的Type 1/TrueType字体,从根源解决问题:
gs -sDEVICE=pdfwrite -dEmbedAllFonts=true -dSubsetFonts=true -dCompressFonts=true -dConvertType3ToType1=true -o output_fixed.pdf input_ghostscript.pdf
核心参数说明:-dConvertType3ToType1=true 强制将Type 3字体转换为Type 1格式,后续文本提取工具可正常识别。
方案2:修改pdfminer.six配置绕过字体检查
通过调整pdfminer.six的内部参数,跳过Type 3字体的可提取性检查,实现文本提取:
from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter from pdfminer.pdfpage import PDFPage from pdfminer.converter import TextConverter from pdfminer.layout import LAParams import io resource_manager = PDFResourceManager() output_buffer = io.StringIO() laparams = LAParams() text_converter = TextConverter(resource_manager, output_buffer, laparams=laparams) interpreter = PDFPageInterpreter(resource_manager, text_converter) # 关键:关闭字体可提取性检查 interpreter._device.check_extractable = False with open("input.pdf", "rb") as pdf_file: for page in PDFPage.get_pages(pdf_file, caching=True, check_extractable=False): interpreter.process_page(page) extracted_text = output_buffer.getvalue() text_converter.close() output_buffer.close()
注意:此方法可能导致部分特殊排版文本提取不准确,但能确保流程不中断。
方案3:调整fitz(PyMuPDF)的文本提取策略
针对fitz提供两种规避方式:
方式A:启用忽略字体的提取模式
import fitz pdf_doc = fitz.open("input.pdf") total_text = "" for page in pdf_doc: # 使用TEXTFLAGS_IGNORE_FONTS跳过字体兼容性检查 page_text = page.get_text("text", flags=fitz.TEXTFLAGS_IGNORE_FONTS) total_text += page_text pdf_doc.close()
方式B:转为图片后OCR提取
如果直接提取失效,可将PDF页面转为图片,通过OCR工具提取文本:
import fitz import pytesseract from PIL import Image pdf_doc = fitz.open("input.pdf") total_text = "" for page_num in range(pdf_doc.page_count): page = pdf_doc.load_page(page_num) pix = page.get_pixmap() img = Image.frombytes("RGB", [pix.width, pix.height], pix.samples) # 需要提前安装Tesseract OCR引擎 page_text = pytesseract.image_to_string(img) total_text += page_text pdf_doc.close()
方案4:使用兼容Type 3字体的第三方工具
如果上述方法均不适用,可使用Poppler工具集中的pdftotext,它对Type 3字体的兼容性更好:
pdftotext input_ghostscript.pdf output_text.txt
Python中调用示例:
import subprocess subprocess.run(["pdftotext", "input.pdf", "output.txt"], check=True)
内容的提问来源于stack exchange,提问作者Abhishek Yadav
相关产品推荐
相关产品推荐

