You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python/R实现同结构发票批量OCR数据提取?

批量处理同结构发票的OCR方案(Python优先)

核心逻辑

因为是同结构发票,无需全页识别,只需定位固定字段的坐标区域提取内容,再结合批量文件遍历+并行处理,就能解决数千张发票的耗时问题,同时兼容你现有的Poppler+Tesseract工具链。


步骤1:批量遍历发票文件

先写脚本遍历指定文件夹下的所有发票(支持PDF/图片格式):

import os

# 发票存放文件夹路径
invoice_dir = "./invoices"
# 支持的文件格式
supported_formats = (".pdf", ".png", ".jpg", ".jpeg")

# 批量获取所有发票文件路径
invoice_files = [
    os.path.join(invoice_dir, f)
    for f in os.listdir(invoice_dir)
    if f.lower().endswith(supported_formats)
]

步骤2:优化单张发票的提取逻辑(针对同结构)

拿一张样本发票,标记好每个字段的坐标(左、上、右、下),后续直接提取对应区域的OCR结果,跳过全页识别:

import pytesseract
from pdf2image import convert_from_path
from PIL import Image, ImageOps

def extract_invoice_data(file_path):
    # 处理PDF文件:转成单页图片(假设发票都是单页)
    if file_path.lower().endswith(".pdf"):
        pages = convert_from_path(file_path, dpi=300)
        img = pages[0]
    else:
        img = Image.open(file_path)
    
    # 预定义同结构发票的字段坐标(需用样本发票调整数值)
    field_coords = {
        "发票号码": (120, 210, 320, 260),
        "合计金额": (520, 410, 720, 460),
        "开票日期": (120, 310, 320, 360)
    }
    
    data = {}
    for field, coords in field_coords.items():
        # 裁剪指定区域
        cropped_img = img.crop(coords)
        # 图像预处理(灰度化+二值化,提升OCR准确率)
        cropped_img = ImageOps.grayscale(cropped_img)
        cropped_img = cropped_img.point(lambda x: 0 if x < 128 else 255, "1")
        # OCR提取文本
        text = pytesseract.image_to_string(cropped_img, lang="chi_sim")
        data[field] = text.strip()
    
    return data

步骤3:并行处理提升效率

用多线程并行处理,替代单线程逐一处理,大幅缩短耗时:

from concurrent.futures import ThreadPoolExecutor
import csv

# 设定并行线程数(根据CPU核心数调整,比如8核设为8)
max_workers = 8

# 批量并行处理所有发票
with ThreadPoolExecutor(max_workers=max_workers) as executor:
    results = list(executor.map(extract_invoice_data, invoice_files))

# 将结果保存为CSV,方便后续统计分析
with open("./发票提取结果.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["发票号码", "合计金额", "开票日期"])
    writer.writeheader()
    writer.writerows(results)

R语言替代方案(可选)

如果偏好R,核心逻辑一致,用pdftools转PDF为图片,tesseract做OCR,furrr实现并行:

library(pdftools)
library(tesseract)
library(furrr)

# 设置并行会话
plan(multisession, workers = 8)

# 遍历发票文件
invoice_dir <- "./invoices"
invoice_files <- list.files(invoice_dir, pattern = "\\.(pdf|png|jpg|jpeg)$", full.names = TRUE)

# 单张发票提取函数
extract_invoice_data <- function(file_path) {
    # 处理PDF
    if(grepl("\\.pdf$", tolower(file_path))) {
        img_path <- pdf_convert(file_path, dpi = 300)[1]
        img <- image_read(img_path)
    } else {
        img <- image_read(file_path)
    }
    
    # 预定义字段坐标(单位:像素)
    field_coords <- list(
        发票号码 = c(120, 210, 320, 260),
        合计金额 = c(520, 410, 720, 460),
        开票日期 = c(120, 310, 320, 360)
    )
    
    data <- list()
    for(field in names(field_coords)) {
        coords <- field_coords[[field]]
        # 裁剪区域
        cropped_img <- image_crop(img, paste0(coords[3]-coords[1], "x", coords[4]-coords[2], "+", coords[1], "+", coords[2]))
        # OCR提取文本
        text <- ocr(cropped_img, engine = tesseract("chi_sim"))
        data[[field]] <- trimws(text)
    }
    
    return(data)
}

# 批量并行处理并保存结果
results <- future_map_dfr(invoice_files, extract_invoice_data)
write.csv(results, "./发票提取结果.csv", row.names = FALSE, fileEncoding = "UTF-8")

额外优化建议

  • 分类模板:如果发票有几种细微差异的结构,可按发票类型分类,每种类型对应一套坐标模板,处理前先判断类型再调用对应模板。
  • 错误校验:对提取的关键字段(如金额)做格式校验,不符合规则的标记出来人工复核。

内容的提问来源于stack exchange,提问作者Gustav Barrows

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 13:35:15