You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdfminer与fitz处理Ghostscript生成PDF时遇Type3字体不支持错误

解决Ghostscript生成PDF的Type 3字体文本提取报错问题

核心问题背景

Ghostscript生成的PDF包含Type 3字体时,pdfminer.six(20221105版本)和fitz(PyMuPDF)默认会触发RuntimeError: pdf device does not support type 3 fonts,导致文本提取流程中断。以下是经过验证的解决和规避方案:


方案1:重新生成PDF时转换Type 3字体

直接通过Ghostscript参数,在生成PDF阶段将Type 3字体转换为工具支持的Type 1/TrueType字体,从根源解决问题:

gs -sDEVICE=pdfwrite -dEmbedAllFonts=true -dSubsetFonts=true -dCompressFonts=true -dConvertType3ToType1=true -o output_fixed.pdf input_ghostscript.pdf

核心参数说明:-dConvertType3ToType1=true 强制将Type 3字体转换为Type 1格式,后续文本提取工具可正常识别。

方案2:修改pdfminer.six配置绕过字体检查

通过调整pdfminer.six的内部参数,跳过Type 3字体的可提取性检查,实现文本提取:

from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter
from pdfminer.pdfpage import PDFPage
from pdfminer.converter import TextConverter
from pdfminer.layout import LAParams
import io

resource_manager = PDFResourceManager()
output_buffer = io.StringIO()
laparams = LAParams()
text_converter = TextConverter(resource_manager, output_buffer, laparams=laparams)
interpreter = PDFPageInterpreter(resource_manager, text_converter)

# 关键:关闭字体可提取性检查
interpreter._device.check_extractable = False

with open("input.pdf", "rb") as pdf_file:
    for page in PDFPage.get_pages(pdf_file, caching=True, check_extractable=False):
        interpreter.process_page(page)

extracted_text = output_buffer.getvalue()
text_converter.close()
output_buffer.close()

注意:此方法可能导致部分特殊排版文本提取不准确,但能确保流程不中断。

方案3:调整fitz(PyMuPDF)的文本提取策略

针对fitz提供两种规避方式:

方式A:启用忽略字体的提取模式

import fitz

pdf_doc = fitz.open("input.pdf")
total_text = ""
for page in pdf_doc:
    # 使用TEXTFLAGS_IGNORE_FONTS跳过字体兼容性检查
    page_text = page.get_text("text", flags=fitz.TEXTFLAGS_IGNORE_FONTS)
    total_text += page_text
pdf_doc.close()

方式B:转为图片后OCR提取

如果直接提取失效,可将PDF页面转为图片,通过OCR工具提取文本:

import fitz
import pytesseract
from PIL import Image

pdf_doc = fitz.open("input.pdf")
total_text = ""
for page_num in range(pdf_doc.page_count):
    page = pdf_doc.load_page(page_num)
    pix = page.get_pixmap()
    img = Image.frombytes("RGB", [pix.width, pix.height], pix.samples)
    # 需要提前安装Tesseract OCR引擎
    page_text = pytesseract.image_to_string(img)
    total_text += page_text
pdf_doc.close()

方案4:使用兼容Type 3字体的第三方工具

如果上述方法均不适用,可使用Poppler工具集中的pdftotext,它对Type 3字体的兼容性更好:

pdftotext input_ghostscript.pdf output_text.txt

Python中调用示例:

import subprocess

subprocess.run(["pdftotext", "input.pdf", "output.txt"], check=True)

内容的提问来源于stack exchange,提问作者Abhishek Yadav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 23:03:32