You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyPDF2读取Roboto字体PDF触发IndexError的解决求助

PyPDF2提取Roboto字体PDF文本触发IndexError的解决方案

问题背景

编写脚本自动提取PDF数据时,使用PyPDF2读取文本出现如下问题:

  • 使用Arial字体的PDF可正常提取文本
  • 使用Roboto字体的PDF触发IndexError: list index out of range

已确认问题根源为Roboto字体,手动将字体改为Arial后恢复正常,但每月需处理数百份此类发票,手动修改不可行。

出错代码片段

import PyPDF2

pdf_roboto = r"C:\Users\Robert.Smyth\Python\test_pdf_roboto.pdf"
pdf_arial = r"C:\Users\Robert.Smyth\Python\test_pdf_arial.pdf"

reader = PyPDF2.PdfFileReader(pdf_roboto)
pageObj = reader.pages[0]
pages_text = pageObj.extractText()

报错信息

---------------------------------------------------------------------------
IndexError                                Traceback (most recent call last)
C:\Users\ROBERT~1.SMY\AppData\Local\Temp/ipykernel_22076/669450932.py in <module>
      1 reader = PyPDF2.PdfFileReader(pdf_roboto)
      2 pageObj = reader.pages[0]
----> 3 pages_text = pageObj.extractText()

~\Anaconda3\lib\site-packages\PyPDF2\_page.py in extractText(self, Tj_sep, TJ_sep)
   1539         """
   1540         deprecate_with_replacement("extractText", "extract_text")
-> 1541         return self.extract_text()
   1542 
   1543     def _get_fonts(self) -> Tuple[Set[str], Set[str]]:

~\Anaconda3\lib\site-packages\PyPDF2\_page.py in extract_text(self, Tj_sep, TJ_sep, orientations, space_width, *args)
   1511             orientations = (orientations,)
   1512 
-> 1513         return self._extract_text(
   1514             self, self.pdf, orientations, space_width, PG.CONTENTS
   1515         )

~\Anaconda3\lib\site-packages\PyPDF2\_page.py in _extract_text(self, obj, pdf, orientations, space_width, content_key)
   1144         if "/Font" in resources_dict:
   1145             for f in cast(DictionaryObject, resources_dict["/Font"]):
-> 1146                 cmaps[f] = build_char_map(f, space_width, obj)
   1147         cmap: Tuple[Union[str, Dict[int, str]], Dict[str, str], str] = (
   1148             "charmap",

~\Anaconda3\lib\site-packages\PyPDF2\_cmap.py in build_char_map(font_name, space_width, obj)
     20     space_code = 32
     21     encoding, space_code = parse_encoding(ft, space_code)
---> 22     map_dict, space_code, int_entry = parse_to_unicode(ft, space_code)
     23 
     24     # encoding can be either a string for decode (on 1,2 or a variable number of bytes) of a char table (for 1 byte only for me)

~\Anaconda3\lib\site-packages\PyPDF2\_cmap.py in parse_to_unicode(ft, space_code)
    187     cm = prepare_cm(ft)
    188     for l in cm.split(b"\n"):
---> 189         process_rg, process_char = process_cm_line(
    190             l.strip(b" "), process_rg, process_char, map_dict, int_entry
    191         )

~\Anaconda3\lib\site-packages\PyPDF2\_cmap.py in process_cm_line(l, process_rg, process_char, map_dict, int_entry)
    247         process_char = False
    248     elif process_rg:
---> 249         parse_bfrange(l, map_dict, int_entry)
    250     elif process_char:
    251         parse_bfchar(l, map_dict, int_entry)

~\Anaconda3\lib\site-packages\PyPDF2\_cmap.py in parse_bfrange(l, map_dict, int_entry)
    256     lst = [x for x in l.split(b" ") if x]
    257     a = int(lst[0], 16)
---> 258     b = int(lst[1], 16)
    259     nbi = len(lst[0])
    260     map_dict[-1] = nbi // 2

IndexError: list index out of range

解决建议

1. 切换到PyMuPDF(fitz)替代PyPDF2

PyMuPDF对复杂字体的兼容性优于PyPDF2,尤其适合Roboto这类非默认系统字体的PDF文本提取。示例代码:

import fitz  # 先安装:pip install pymupdf

pdf_roboto = r"C:\Users\Robert.Smyth\Python\test_pdf_roboto.pdf"
doc = fitz.open(pdf_roboto)
page = doc[0]
pages_text = page.get_text()
print(pages_text)

2. 升级PyPDF2到最新版本

该IndexError可能是PyPDF2旧版本的已知bug,更新到最新版可能修复:

pip install --upgrade PyPDF2

3. 预处理PDF转换为文本友好格式

如果上述方法无效,可先用工具将PDF转换为标准化格式后再提取文本。例如使用Ghostscript转换(需先安装Ghostscript):

gswin64c -sDEVICE=pdfwrite -dCompatibilityLevel=1.4 -dPDFSETTINGS=/prepress -dNOPAUSE -dBATCH -sOutputFile=output.pdf input.pdf

转换后的PDF再用PyPDF2提取文本。

4. 临时补丁修复PyPDF2的cmap解析逻辑

若必须使用PyPDF2,可修改_cmap.py中的parse_bfrange函数,增加列表长度检查避免索引越界:
找到parse_bfrange函数,修改为:

def parse_bfrange(l: bytes, map_dict: Dict[int, str], int_entry: bool) -> None:
    lst = [x for x in l.split(b" ") if x]
    if len(lst) < 2:
        return  # 跳过无效的bfrange行
    a = int(lst[0], 16)
    b = int(lst[1], 16)
    nbi = len(lst[0])
    map_dict[-1] = nbi // 2
    # 剩余逻辑保持不变

注意:修改库文件需谨慎,后续更新PyPDF2后补丁会被覆盖。

内容的提问来源于stack exchange,提问作者Robert Smyth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 04:10:32