PyPDF2读取带数字签名PDF报IndexError,求跳过签名字段方案
提取PDF表单字段时跳过签名字段解决IndexError问题
问题
我是Python和PyPDF2新手,正在尝试提取PDF表单中的所有字段并存入DataFrame,最终要批量处理数千份同结构的PDF表单。代码在无数字签名的PDF上运行正常,但读取带数字签名的PDF时触发IndexError。由于不需要签名字段,希望跳过该字段但不知如何实现。
原代码
import os import PyPDF2 as pypdf import pandas as pd directory = 'files' for filename in os.listdir(directory): f = os.path.join(directory, filename) if os.path.isfile(f): print(f) pdf=pypdf.PdfFileReader(f, strict= False) print(pdf) #information = pdf.getFormTextFields() information = pdf.getFields() print(information) output = pd.DataFrame([information]) df = pd.concat([df, output], ignore_index=True)
错误信息
Traceback (most recent call last): File "/workspace/app.py", line 77, in <module> information = pdf.getFields() File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_reader.py", line 526, in getFields return self.get_fields(tree, retval, fileobj) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_reader.py", line 510, in get_fields self._build_field(field, retval, fileobj, field_attributes) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_reader.py", line 535, in _build_field self._check_kids(field, retval, fileobj) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_reader.py", line 555, in _check_kids self.get_fields(kid.get_object(), retval, fileobj) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_reader.py", line 499, in get_fields self._check_kids(tree, retval, fileobj) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_reader.py", line 555, in _check_kids self.get_fields(kid.get_object(), retval, fileobj) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_reader.py", line 503, in get_fields self._build_field(tree, retval, fileobj, field_attributes) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_reader.py", line 547, in _build_field retval[key] = Field(field) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/generic.py", line 1626, in __init__ self[NameObject(attr)] = data[attr] File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/generic.py", line 679, in __getitem__ return dict.__getitem__(self, key).get_object() File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/generic.py", line 251, in get_object obj = self.pdf.get_object(self) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_reader.py", line 1167, in get_object retval, indirect_reference.idnum, indirect_reference.generation File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_encryption.py", line 741, in decrypt_object return cf.decrypt_object(obj) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_encryption.py", line 182, in decrypt_object obj[dictkey] = self.decrypt_object(value) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_encryption.py", line 185, in decrypt_object obj[i] = self.decrypt_object(obj[i]) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_encryption.py", line 182, in decrypt_object obj[dictkey] = self.decrypt_object(value) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_encryption.py", line 176, in decrypt_object data = self.strCrypt.decrypt(obj.original_bytes) File "/app/.heroku/python/lib/python3.7/site-packages/PyPDF2/_encryption.py", line 88, in decrypt return d[: -d[-1]] IndexError: index out of range
解决方案
方法1:捕获异常+过滤签名字段
用try-except包裹字段读取操作,遇到IndexError时降级到纯文本字段读取;读取成功后手动过滤掉类型为/Sig的签名字段:
import os import PyPDF2 as pypdf import pandas as pd directory = 'files' df = pd.DataFrame() # 初始化空DataFrame,避免未定义报错 for filename in os.listdir(directory): f = os.path.join(directory, filename) # 只处理PDF文件,避免非PDF干扰 if os.path.isfile(f) and f.lower().endswith('.pdf'): print(f"Processing: {f}") pdf = pypdf.PdfFileReader(f, strict=False) information = {} try: # 尝试读取所有字段 information = pdf.getFields() # 过滤签名字段:签名字段的类型标识为'/Sig' information = {k: v for k, v in information.items() if v.get('/FT') != '/Sig'} except IndexError: # 读取失败时,改用getFormTextFields()提取纯文本字段 print(f"Error reading {f}, falling back to text fields only") information = pdf.getFormTextFields() # 将有效字段加入DataFrame if information: output = pd.DataFrame([information]) df = pd.concat([df, output], ignore_index=True) # 保存结果到CSV df.to_csv('pdf_form_fields.csv', index=False)
方法2:手动遍历表单结构跳过签名字段
直接遍历PDF的AcroForm结构,从根源跳过签名字段,避免触发读取异常:
import os import PyPDF2 as pypdf import pandas as pd from PyPDF2.generic import NameObject directory = 'files' df = pd.DataFrame() for filename in os.listdir(directory): f = os.path.join(directory, filename) if os.path.isfile(f) and f.lower().endswith('.pdf'): print(f"Processing: {f}") pdf = pypdf.PdfFileReader(f, strict=False) information = {} acro_form = pdf.getAcroForm() if acro_form is not None: # 递归遍历所有字段节点 def traverse_fields(node, field_dict): if '/Kids' in node: for kid in node['/Kids']: traverse_fields(kid.get_object(), field_dict) elif '/FT' in node: # 跳过类型为'/Sig'的签名字段 if node['/FT'] != NameObject('/Sig'): # 处理字段名称和值 field_name = node.get('/T', '') if isinstance(field_name, bytes): field_name = field_name.decode('utf-8') field_value = node.get('/V') if field_value is not None: if isinstance(field_value, bytes): field_value = field_value.decode('utf-8') else: field_value = str(field_value) field_dict[field_name] = field_value traverse_fields(acro_form, information) # 加入DataFrame if information: output = pd.DataFrame([information]) df = pd.concat([df, output], ignore_index=True) df.to_csv('pdf_form_fields.csv', index=False)
关键说明
- 签名字段的标准类型标识为
/Sig,通过判断字段的/FT属性即可过滤。 - 方法1实现简单,优先保证兼容性,适合大多数场景。
- 方法2直接操作PDF底层结构,避免读取签名字段时触发异常,适合报错频繁的批量处理场景。
内容的提问来源于stack exchange,提问作者ItzGravy314
相关产品推荐
相关产品推荐

