You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取PDF数字签名信息?现有代码无法识别签名

解决PDF数字签名信息提取问题

原代码无法识别PDF数字签名的核心问题在于:

  • 错误假设签名存储格式:PDF中的数字签名是以二进制形式嵌入在PDF结构的签名字典中,不会包含-----BEGIN PKCS7-----这类明文标记,正则匹配完全无法定位签名数据。
  • 标准库支持有限:Python的ssl.PKCS7模块对签名解析的能力不足,难以准确提取所需的证书字段。

改进方案

使用专门处理PDF的PyMuPDF(fitz)库提取签名数据,结合cryptography库解析PKCS7格式的签名信息,可稳定提取目标字段。

1. 安装依赖

pip install pymupdf cryptography

2. 完整代码

import os
import glob
from cryptography import x509
from cryptography.hazmat.backends import default_backend
from cryptography.hazmat.primitives.serialization import pkcs7
import fitz  # PyMuPDF

def extract_signature_info(pdf_path):
    doc = fitz.open(pdf_path)
    # 遍历PDF中所有签名
    for sig in doc.signatures:
        if not sig.sigbytes:
            continue
        
        # 解析PKCS7签名数据,兼容PEM和DER编码
        try:
            if b"BEGIN CERTIFICATE" in sig.sigbytes:
                cert = x509.load_pem_x509_certificate(sig.sigbytes, default_backend())
            elif b"BEGIN PKCS7" in sig.sigbytes:
                certs = pkcs7.load_pem_pkcs7_certificates(sig.sigbytes, default_backend())
                cert = certs[0]
            else:
                try:
                    certs = pkcs7.load_der_pkcs7_certificates(sig.sigbytes, default_backend())
                    cert = certs[0]
                except:
                    cert = x509.load_der_x509_certificate(sig.sigbytes, default_backend())
        except Exception as e:
            print(f"文件 {pdf_path} 解析失败:{str(e)}")
            continue
        
        # 提取签名者姓名(Common Name)
        try:
            common_name = cert.subject.get_attributes_for_oid(x509.NameOID.COMMON_NAME)[0].value
        except IndexError:
            common_name = "未找到"
        
        # 提取巴西CPF(优先从serialNumber字段获取,清理格式)
        cpf = None
        try:
            serial_num = cert.subject.get_attributes_for_oid(x509.NameOID.SERIAL_NUMBER)[0].value
            cpf = ''.join([c for c in serial_num if c.isdigit()])
            # 验证CPF长度(可选)
            if len(cpf) != 11:
                cpf = None
        except IndexError:
            pass
        
        # 提取签发与过期日期
        issuance_date = cert.not_valid_before.strftime("%Y-%m-%d %H:%M:%S")
        expiration_date = cert.not_valid_after.strftime("%Y-%m-%d %H:%M:%S")
        
        # 输出结果
        print(f"文件:{os.path.basename(pdf_path)}")
        print(f"签名者姓名:{common_name}")
        print(f"CPF:{cpf if cpf else '未找到'}")
        print(f"签发日期:{issuance_date}")
        print(f"过期日期:{expiration_date}")
        print("-" * 50)

# 遍历目标目录下的所有PDF
pdf_dir = r'C:\Users\est.matheus.ro\Desktop\Certificados\push\pdfs'
for filepath in glob.glob(os.path.join(pdf_dir, '*.pdf')):
    extract_signature_info(filepath)

关键说明

  • PyMuPDF自动识别PDF的签名字典,直接提取签名二进制数据,无需手动解析PDF结构。
  • cryptography库支持多种编码格式的PKCS7解析,覆盖不同证书颁发机构的签名格式。
  • 巴西CPF通常存储在证书的serialNumber字段中,代码会自动清理非数字字符;若你的证书CPF存储在其他OID字段,可通过openssl x509 -in 证书文件.pem -text -noout查看详情后调整代码。
  • 代码支持遍历PDF中的多个签名,逐一输出信息。

内容的提问来源于stack exchange,提问作者Matheyc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 09:28:09