You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdfplumber提取PDF文本时出现(cid:x)问题的解决方法

解决PDFplumber提取文本出现(cid:x)的问题

出现大量(cid:x)是因为目标PDF使用了CID字体,字符未映射到标准Unicode编码,导致pdfplumber默认的文本提取无法解析这些字符。以下是两种可行的解决方法:

方法一:调整PDFplumber提取参数

尝试启用布局分析模式,该模式能更好地处理非标准字体的文本排版:

with pdfplumber.open(fatura) as pdf:
    fatura_individual = ''
    for pagina in pdf.pages:
        # 启用布局分析,优化非标准字体的文本提取
        text = pagina.extract_text(layout=True)
        if text:
            fatura_individual += text

方法二:结合OCR工具提取文本

如果调整参数后仍无效,说明PDF的字体编码完全缺失(比如扫描生成的PDF),此时需要用OCR基于图像识别文本:

  1. 先安装依赖:
pip install pytesseract pdf2image

注意:需在本地安装Tesseract OCR引擎,并根据文档语言配置对应的语言包(比如葡萄牙语包por、中文包chi_sim)。

  1. 提取代码:
import pytesseract
from pdf2image import convert_from_path

# 将PDF页面转为图片
pages = convert_from_path(fatura)
fatura_individual = ''

for page in pages:
    # 执行OCR提取文本,lang参数指定文档语言
    text = pytesseract.image_to_string(page, lang='por')
    fatura_individual += text

内容的提问来源于stack exchange,提问作者foliveir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 08:35:35