PyPDF2读取Roboto字体PDF触发IndexError的解决求助
PyPDF2提取Roboto字体PDF文本触发IndexError的解决方案
问题背景
编写脚本自动提取PDF数据时,使用PyPDF2读取文本出现如下问题:
- 使用Arial字体的PDF可正常提取文本
- 使用Roboto字体的PDF触发
IndexError: list index out of range
已确认问题根源为Roboto字体,手动将字体改为Arial后恢复正常,但每月需处理数百份此类发票,手动修改不可行。
出错代码片段
import PyPDF2 pdf_roboto = r"C:\Users\Robert.Smyth\Python\test_pdf_roboto.pdf" pdf_arial = r"C:\Users\Robert.Smyth\Python\test_pdf_arial.pdf" reader = PyPDF2.PdfFileReader(pdf_roboto) pageObj = reader.pages[0] pages_text = pageObj.extractText()
报错信息
--------------------------------------------------------------------------- IndexError Traceback (most recent call last) C:\Users\ROBERT~1.SMY\AppData\Local\Temp/ipykernel_22076/669450932.py in <module> 1 reader = PyPDF2.PdfFileReader(pdf_roboto) 2 pageObj = reader.pages[0] ----> 3 pages_text = pageObj.extractText() ~\Anaconda3\lib\site-packages\PyPDF2\_page.py in extractText(self, Tj_sep, TJ_sep) 1539 """ 1540 deprecate_with_replacement("extractText", "extract_text") -> 1541 return self.extract_text() 1542 1543 def _get_fonts(self) -> Tuple[Set[str], Set[str]]: ~\Anaconda3\lib\site-packages\PyPDF2\_page.py in extract_text(self, Tj_sep, TJ_sep, orientations, space_width, *args) 1511 orientations = (orientations,) 1512 -> 1513 return self._extract_text( 1514 self, self.pdf, orientations, space_width, PG.CONTENTS 1515 ) ~\Anaconda3\lib\site-packages\PyPDF2\_page.py in _extract_text(self, obj, pdf, orientations, space_width, content_key) 1144 if "/Font" in resources_dict: 1145 for f in cast(DictionaryObject, resources_dict["/Font"]): -> 1146 cmaps[f] = build_char_map(f, space_width, obj) 1147 cmap: Tuple[Union[str, Dict[int, str]], Dict[str, str], str] = ( 1148 "charmap", ~\Anaconda3\lib\site-packages\PyPDF2\_cmap.py in build_char_map(font_name, space_width, obj) 20 space_code = 32 21 encoding, space_code = parse_encoding(ft, space_code) ---> 22 map_dict, space_code, int_entry = parse_to_unicode(ft, space_code) 23 24 # encoding can be either a string for decode (on 1,2 or a variable number of bytes) of a char table (for 1 byte only for me) ~\Anaconda3\lib\site-packages\PyPDF2\_cmap.py in parse_to_unicode(ft, space_code) 187 cm = prepare_cm(ft) 188 for l in cm.split(b"\n"): ---> 189 process_rg, process_char = process_cm_line( 190 l.strip(b" "), process_rg, process_char, map_dict, int_entry 191 ) ~\Anaconda3\lib\site-packages\PyPDF2\_cmap.py in process_cm_line(l, process_rg, process_char, map_dict, int_entry) 247 process_char = False 248 elif process_rg: ---> 249 parse_bfrange(l, map_dict, int_entry) 250 elif process_char: 251 parse_bfchar(l, map_dict, int_entry) ~\Anaconda3\lib\site-packages\PyPDF2\_cmap.py in parse_bfrange(l, map_dict, int_entry) 256 lst = [x for x in l.split(b" ") if x] 257 a = int(lst[0], 16) ---> 258 b = int(lst[1], 16) 259 nbi = len(lst[0]) 260 map_dict[-1] = nbi // 2 IndexError: list index out of range
解决建议
1. 切换到PyMuPDF(fitz)替代PyPDF2
PyMuPDF对复杂字体的兼容性优于PyPDF2,尤其适合Roboto这类非默认系统字体的PDF文本提取。示例代码:
import fitz # 先安装:pip install pymupdf pdf_roboto = r"C:\Users\Robert.Smyth\Python\test_pdf_roboto.pdf" doc = fitz.open(pdf_roboto) page = doc[0] pages_text = page.get_text() print(pages_text)
2. 升级PyPDF2到最新版本
该IndexError可能是PyPDF2旧版本的已知bug,更新到最新版可能修复:
pip install --upgrade PyPDF2
3. 预处理PDF转换为文本友好格式
如果上述方法无效,可先用工具将PDF转换为标准化格式后再提取文本。例如使用Ghostscript转换(需先安装Ghostscript):
gswin64c -sDEVICE=pdfwrite -dCompatibilityLevel=1.4 -dPDFSETTINGS=/prepress -dNOPAUSE -dBATCH -sOutputFile=output.pdf input.pdf
转换后的PDF再用PyPDF2提取文本。
4. 临时补丁修复PyPDF2的cmap解析逻辑
若必须使用PyPDF2,可修改_cmap.py中的parse_bfrange函数,增加列表长度检查避免索引越界:
找到parse_bfrange函数,修改为:
def parse_bfrange(l: bytes, map_dict: Dict[int, str], int_entry: bool) -> None: lst = [x for x in l.split(b" ") if x] if len(lst) < 2: return # 跳过无效的bfrange行 a = int(lst[0], 16) b = int(lst[1], 16) nbi = len(lst[0]) map_dict[-1] = nbi // 2 # 剩余逻辑保持不变
注意:修改库文件需谨慎,后续更新PyPDF2后补丁会被覆盖。
内容的提问来源于stack exchange,提问作者Robert Smyth
相关产品推荐
相关产品推荐

