You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现基于图像的多页PDF中EPC相关行及后续行提取问题

提取图像型PDF中EPC行及对应下一行的解决方案

我有一个多页的图像型PDF,需要提取包含ENERGY PERFORMANCE CERTIFICATE的行,以及它的下一行。示例如下:

ENERGY PERFORMANCE CERTIFICATE
D(139)

我尝试了以下Python代码,但未能正确实现需求:

import os    
from PIL import Image
import pytesseract
from pdf2image import convert_from_path
import re

poppler_path = "C:/Users/poddaral/temp/poppler-0.68.0/bin"
pytesseract.pytesseract.tesseract_cmd = r"C:/Users/poddaral/temp/tesseract.exe"

pdf_path = "C:/Users/poddaral/temp/4-6 Etloe Road, Westbury Park, Bristol, BS6 7PF.pdf"

images = convert_from_path(pdf_path=pdf_path, poppler_path=poppler_path)

for count, img in enumerate(images):
  img_name = f"page_{count}.png"  
  img.save(img_name, "PNG")

png_files = [f for f in os.listdir(".") if f.endswith(".png")]

for png_file in png_files:
  extracted_text = pytesseract.image_to_string(Image.open(png_file))
  print(extracted_text)

pattern = re.compile('ENERGY PERFORMANCE CERTIFICATE')

def find_following_line(extracted_text):
    lines = extracted_text.splitlines()
    for i, line in enumerate(lines):
        if re.search(pattern, line):
            return lines[i+2]

print(find_following_line(extracted_text))

问题分析

你的代码存在几个关键问题:

  • extracted_text会被循环覆盖,最终只保留最后一页的文本,前面页面的匹配内容直接丢失
  • 查找下一行时错误使用i+2,目标内容是匹配行的直接下一行,应该用i+1
  • 没有处理边界情况:如果匹配行在页面最后一行,i+1会触发索引越界错误
  • 仅处理了最后一页的内容,没有遍历所有页面的匹配结果

修正后的代码

import pytesseract
from pdf2image import convert_from_path
import re

# 配置工具路径
poppler_path = "C:/Users/poddaral/temp/poppler-0.68.0/bin"
pytesseract.pytesseract.tesseract_cmd = r"C:/Users/poddaral/temp/tesseract.exe"
pdf_path = "C:/Users/poddaral/temp/4-6 Etloe Road, Westbury Park, Bristol, BS6 7PF.pdf"

# 编译正则表达式,支持忽略大小写匹配(可按需移除re.IGNORECASE)
pattern = re.compile(r'ENERGY PERFORMANCE CERTIFICATE', re.IGNORECASE)

def extract_epc_content(text):
    # 清理空白行,只保留有效文本行
    lines = [line.strip() for line in text.splitlines() if line.strip()]
    results = []
    for i, line in enumerate(lines):
        if pattern.search(line):
            # 检查下一行是否存在,避免索引越界
            if i + 1 < len(lines):
                results.append({
                    'epc_line': line,
                    'next_line': lines[i+1]
                })
            else:
                results.append({
                    'epc_line': line,
                    'next_line': '无后续内容'
                })
    return results

# 转换PDF为图像并逐页处理
images = convert_from_path(pdf_path=pdf_path, poppler_path=poppler_path)

for page_num, img in enumerate(images, start=1):
    print(f"=== 第 {page_num} 页提取结果 ===")
    extracted_text = pytesseract.image_to_string(img)
    epc_results = extract_epc_content(extracted_text)
    if epc_results:
        for res in epc_results:
            print(res['epc_line'])
            print(res['next_line'])
            print("---")
    else:
        print("未找到EPC相关内容")

修正说明

  1. 直接处理PDF转换后的图像,无需保存为PNG文件,节省磁盘资源
  2. 遍历所有页面,不会遗漏任何一页的匹配内容
  3. 处理边界情况,避免索引越界错误
  4. 支持忽略大小写匹配,提升识别容错率
  5. 清理空白行,减少无效文本干扰,提升匹配准确性
  6. 返回所有匹配结果,而非仅第一个匹配项

内容的提问来源于stack exchange,提问作者Alok

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 09:25:46