如何将OCR提取的文本添加至pypdf PageObject并覆写PDF?
解决方案:为PDF添加OCR文本层并保存
核心结论
pypdf从PdfReader获取的PageObject是只读实例,无法直接修改并覆写原PDF。正确的做法是创建新PDF,复制原页面内容后,将OCR提取的文本作为隐藏文本层添加进去——既保留原PDF的视觉效果,后续使用extract_text()时又能直接读取到OCR文本。
实现代码
from pypdf import PdfReader, PdfWriter from pypdf.generic import NameObject, DictionaryObject, TextStringObject def add_hidden_ocr_text(page, ocr_text): # 初始化页面字体资源 if "/Resources" not in page: page[NameObject("/Resources")] = DictionaryObject() if "/Font" not in page["/Resources"]: page["/Resources"][NameObject("/Font")] = DictionaryObject() # 添加基础Helvetica字体 page["/Resources"]["/Font"][NameObject("/F1")] = DictionaryObject({ NameObject("/Type"): NameObject("/Font"), NameObject("/Subtype"): NameObject("/Type1"), NameObject("/BaseFont"): NameObject("/Helvetica"), NameObject("/Encoding"): NameObject("/WinAnsiEncoding") }) # 构建隐藏文本的PDF指令:设置文本颜色为白色(与背景融合),转义特殊字符 escaped_text = ocr_text.replace("(", "\\(").replace(")", "\\)") text_content = f""" BT /F1 10 Tf 1 1 1 rg % 文本颜色设为白色,与常见背景融合 50 50 Td % 文本起始位置,可按需调整 ({escaped_text}) Tj ET """ # 合并原有页面内容与新文本内容 if page.get("/Contents"): existing_content = page["/Contents"].decode() new_content = existing_content + text_content else: new_content = text_content page[NameObject("/Contents")] = TextStringObject(new_content) # 主处理流程 input_pdf = "your_input.pdf" output_pdf = "output_with_ocr.pdf" reader = PdfReader(input_pdf) writer = PdfWriter() for page in reader.pages: # 检查页面是否已有可提取文本 raw_text = page.extract_text() or "" if not raw_text.strip(): # 对页面图片执行OCR ocr_text = "" for img in page.images: ocr_text += extract_text_from_image(img) raw_text = ocr_text # 复制原页面到输出PDF并添加隐藏文本层 new_page = writer.add_page(page) add_hidden_ocr_text(new_page, raw_text) # 保存带OCR文本的PDF with open(output_pdf, "wb") as out_file: writer.write(out_file)
优化建议
- 精准文本定位:如果需要文本位置与图片中的文字对应,可以使用Tesseract的HOCR输出格式,解析每个文字的坐标后生成对应位置的PDF文本指令,避免文本堆积导致提取顺序混乱。
- 保留交互元素:上述方法能完整保留原PDF的表单、链接等交互元素;若直接用OCR工具生成新PDF,可能会丢失这些元素。
- 加密PDF处理:若原PDF有密码保护,需先调用
reader.decrypt("password")解密后再处理。 - 批量处理:可封装为函数,批量处理文件夹内的所有扫描版PDF。
内容的提问来源于stack exchange,提问作者Jonas Neubürger
相关产品推荐
相关产品推荐

