You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将OCR提取的文本添加至pypdf PageObject并覆写PDF?

解决方案:为PDF添加OCR文本层并保存

核心结论

pypdf从PdfReader获取的PageObject是只读实例,无法直接修改并覆写原PDF。正确的做法是创建新PDF,复制原页面内容后,将OCR提取的文本作为隐藏文本层添加进去——既保留原PDF的视觉效果,后续使用extract_text()时又能直接读取到OCR文本。

实现代码

from pypdf import PdfReader, PdfWriter
from pypdf.generic import NameObject, DictionaryObject, TextStringObject

def add_hidden_ocr_text(page, ocr_text):
    # 初始化页面字体资源
    if "/Resources" not in page:
        page[NameObject("/Resources")] = DictionaryObject()
    if "/Font" not in page["/Resources"]:
        page["/Resources"][NameObject("/Font")] = DictionaryObject()
    
    # 添加基础Helvetica字体
    page["/Resources"]["/Font"][NameObject("/F1")] = DictionaryObject({
        NameObject("/Type"): NameObject("/Font"),
        NameObject("/Subtype"): NameObject("/Type1"),
        NameObject("/BaseFont"): NameObject("/Helvetica"),
        NameObject("/Encoding"): NameObject("/WinAnsiEncoding")
    })

    # 构建隐藏文本的PDF指令:设置文本颜色为白色(与背景融合),转义特殊字符
    escaped_text = ocr_text.replace("(", "\\(").replace(")", "\\)")
    text_content = f"""
    BT
    /F1 10 Tf
    1 1 1 rg  % 文本颜色设为白色,与常见背景融合
    50 50 Td  % 文本起始位置,可按需调整
    ({escaped_text}) Tj
    ET
    """

    # 合并原有页面内容与新文本内容
    if page.get("/Contents"):
        existing_content = page["/Contents"].decode()
        new_content = existing_content + text_content
    else:
        new_content = text_content
    
    page[NameObject("/Contents")] = TextStringObject(new_content)

# 主处理流程
input_pdf = "your_input.pdf"
output_pdf = "output_with_ocr.pdf"

reader = PdfReader(input_pdf)
writer = PdfWriter()

for page in reader.pages:
    # 检查页面是否已有可提取文本
    raw_text = page.extract_text() or ""
    if not raw_text.strip():
        # 对页面图片执行OCR
        ocr_text = ""
        for img in page.images:
            ocr_text += extract_text_from_image(img)
        raw_text = ocr_text
    
    # 复制原页面到输出PDF并添加隐藏文本层
    new_page = writer.add_page(page)
    add_hidden_ocr_text(new_page, raw_text)

# 保存带OCR文本的PDF
with open(output_pdf, "wb") as out_file:
    writer.write(out_file)

优化建议

  1. 精准文本定位:如果需要文本位置与图片中的文字对应,可以使用Tesseract的HOCR输出格式,解析每个文字的坐标后生成对应位置的PDF文本指令,避免文本堆积导致提取顺序混乱。
  2. 保留交互元素:上述方法能完整保留原PDF的表单、链接等交互元素;若直接用OCR工具生成新PDF,可能会丢失这些元素。
  3. 加密PDF处理:若原PDF有密码保护,需先调用reader.decrypt("password")解密后再处理。
  4. 批量处理:可封装为函数,批量处理文件夹内的所有扫描版PDF。

内容的提问来源于stack exchange,提问作者Jonas Neubürger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 07:52:44