Python实现基于图像的多页PDF中EPC相关行及后续行提取问题
提取图像型PDF中EPC行及对应下一行的解决方案
我有一个多页的图像型PDF,需要提取包含ENERGY PERFORMANCE CERTIFICATE的行,以及它的下一行。示例如下:
ENERGY PERFORMANCE CERTIFICATE
D(139)
我尝试了以下Python代码,但未能正确实现需求:
import os from PIL import Image import pytesseract from pdf2image import convert_from_path import re poppler_path = "C:/Users/poddaral/temp/poppler-0.68.0/bin" pytesseract.pytesseract.tesseract_cmd = r"C:/Users/poddaral/temp/tesseract.exe" pdf_path = "C:/Users/poddaral/temp/4-6 Etloe Road, Westbury Park, Bristol, BS6 7PF.pdf" images = convert_from_path(pdf_path=pdf_path, poppler_path=poppler_path) for count, img in enumerate(images): img_name = f"page_{count}.png" img.save(img_name, "PNG") png_files = [f for f in os.listdir(".") if f.endswith(".png")] for png_file in png_files: extracted_text = pytesseract.image_to_string(Image.open(png_file)) print(extracted_text) pattern = re.compile('ENERGY PERFORMANCE CERTIFICATE') def find_following_line(extracted_text): lines = extracted_text.splitlines() for i, line in enumerate(lines): if re.search(pattern, line): return lines[i+2] print(find_following_line(extracted_text))
问题分析
你的代码存在几个关键问题:
extracted_text会被循环覆盖,最终只保留最后一页的文本,前面页面的匹配内容直接丢失- 查找下一行时错误使用
i+2,目标内容是匹配行的直接下一行,应该用i+1 - 没有处理边界情况:如果匹配行在页面最后一行,
i+1会触发索引越界错误 - 仅处理了最后一页的内容,没有遍历所有页面的匹配结果
修正后的代码
import pytesseract from pdf2image import convert_from_path import re # 配置工具路径 poppler_path = "C:/Users/poddaral/temp/poppler-0.68.0/bin" pytesseract.pytesseract.tesseract_cmd = r"C:/Users/poddaral/temp/tesseract.exe" pdf_path = "C:/Users/poddaral/temp/4-6 Etloe Road, Westbury Park, Bristol, BS6 7PF.pdf" # 编译正则表达式,支持忽略大小写匹配(可按需移除re.IGNORECASE) pattern = re.compile(r'ENERGY PERFORMANCE CERTIFICATE', re.IGNORECASE) def extract_epc_content(text): # 清理空白行,只保留有效文本行 lines = [line.strip() for line in text.splitlines() if line.strip()] results = [] for i, line in enumerate(lines): if pattern.search(line): # 检查下一行是否存在,避免索引越界 if i + 1 < len(lines): results.append({ 'epc_line': line, 'next_line': lines[i+1] }) else: results.append({ 'epc_line': line, 'next_line': '无后续内容' }) return results # 转换PDF为图像并逐页处理 images = convert_from_path(pdf_path=pdf_path, poppler_path=poppler_path) for page_num, img in enumerate(images, start=1): print(f"=== 第 {page_num} 页提取结果 ===") extracted_text = pytesseract.image_to_string(img) epc_results = extract_epc_content(extracted_text) if epc_results: for res in epc_results: print(res['epc_line']) print(res['next_line']) print("---") else: print("未找到EPC相关内容")
修正说明
- 直接处理PDF转换后的图像,无需保存为PNG文件,节省磁盘资源
- 遍历所有页面,不会遗漏任何一页的匹配内容
- 处理边界情况,避免索引越界错误
- 支持忽略大小写匹配,提升识别容错率
- 清理空白行,减少无效文本干扰,提升匹配准确性
- 返回所有匹配结果,而非仅第一个匹配项
内容的提问来源于stack exchange,提问作者Alok
相关产品推荐
相关产品推荐

