如何用Python/R实现同结构发票批量OCR数据提取?
批量处理同结构发票的OCR方案(Python优先)
核心逻辑
因为是同结构发票,无需全页识别,只需定位固定字段的坐标区域提取内容,再结合批量文件遍历+并行处理,就能解决数千张发票的耗时问题,同时兼容你现有的Poppler+Tesseract工具链。
步骤1:批量遍历发票文件
先写脚本遍历指定文件夹下的所有发票(支持PDF/图片格式):
import os # 发票存放文件夹路径 invoice_dir = "./invoices" # 支持的文件格式 supported_formats = (".pdf", ".png", ".jpg", ".jpeg") # 批量获取所有发票文件路径 invoice_files = [ os.path.join(invoice_dir, f) for f in os.listdir(invoice_dir) if f.lower().endswith(supported_formats) ]
步骤2:优化单张发票的提取逻辑(针对同结构)
拿一张样本发票,标记好每个字段的坐标(左、上、右、下),后续直接提取对应区域的OCR结果,跳过全页识别:
import pytesseract from pdf2image import convert_from_path from PIL import Image, ImageOps def extract_invoice_data(file_path): # 处理PDF文件:转成单页图片(假设发票都是单页) if file_path.lower().endswith(".pdf"): pages = convert_from_path(file_path, dpi=300) img = pages[0] else: img = Image.open(file_path) # 预定义同结构发票的字段坐标(需用样本发票调整数值) field_coords = { "发票号码": (120, 210, 320, 260), "合计金额": (520, 410, 720, 460), "开票日期": (120, 310, 320, 360) } data = {} for field, coords in field_coords.items(): # 裁剪指定区域 cropped_img = img.crop(coords) # 图像预处理(灰度化+二值化,提升OCR准确率) cropped_img = ImageOps.grayscale(cropped_img) cropped_img = cropped_img.point(lambda x: 0 if x < 128 else 255, "1") # OCR提取文本 text = pytesseract.image_to_string(cropped_img, lang="chi_sim") data[field] = text.strip() return data
步骤3:并行处理提升效率
用多线程并行处理,替代单线程逐一处理,大幅缩短耗时:
from concurrent.futures import ThreadPoolExecutor import csv # 设定并行线程数(根据CPU核心数调整,比如8核设为8) max_workers = 8 # 批量并行处理所有发票 with ThreadPoolExecutor(max_workers=max_workers) as executor: results = list(executor.map(extract_invoice_data, invoice_files)) # 将结果保存为CSV,方便后续统计分析 with open("./发票提取结果.csv", "w", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=["发票号码", "合计金额", "开票日期"]) writer.writeheader() writer.writerows(results)
R语言替代方案(可选)
如果偏好R,核心逻辑一致,用pdftools转PDF为图片,tesseract做OCR,furrr实现并行:
library(pdftools) library(tesseract) library(furrr) # 设置并行会话 plan(multisession, workers = 8) # 遍历发票文件 invoice_dir <- "./invoices" invoice_files <- list.files(invoice_dir, pattern = "\\.(pdf|png|jpg|jpeg)$", full.names = TRUE) # 单张发票提取函数 extract_invoice_data <- function(file_path) { # 处理PDF if(grepl("\\.pdf$", tolower(file_path))) { img_path <- pdf_convert(file_path, dpi = 300)[1] img <- image_read(img_path) } else { img <- image_read(file_path) } # 预定义字段坐标(单位:像素) field_coords <- list( 发票号码 = c(120, 210, 320, 260), 合计金额 = c(520, 410, 720, 460), 开票日期 = c(120, 310, 320, 360) ) data <- list() for(field in names(field_coords)) { coords <- field_coords[[field]] # 裁剪区域 cropped_img <- image_crop(img, paste0(coords[3]-coords[1], "x", coords[4]-coords[2], "+", coords[1], "+", coords[2])) # OCR提取文本 text <- ocr(cropped_img, engine = tesseract("chi_sim")) data[[field]] <- trimws(text) } return(data) } # 批量并行处理并保存结果 results <- future_map_dfr(invoice_files, extract_invoice_data) write.csv(results, "./发票提取结果.csv", row.names = FALSE, fileEncoding = "UTF-8")
额外优化建议
- 分类模板:如果发票有几种细微差异的结构,可按发票类型分类,每种类型对应一套坐标模板,处理前先判断类型再调用对应模板。
- 错误校验:对提取的关键字段(如金额)做格式校验,不符合规则的标记出来人工复核。
内容的提问来源于stack exchange,提问作者Gustav Barrows
相关产品推荐
相关产品推荐

