You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python Tesseract优化扫描单据表格数据提取的技术问询

针对扫描索赔报告OCR处理的优化方案

1. 针对固定格式列优化识别精度

强化图像预处理

在现有预处理流程基础上,增加二值化处理,强化文字与背景的边界,减少识别干扰:

# 转灰度后添加二值化步骤
image = image.convert('L')
# 自定义阈值(可根据图片清晰度调整,建议160-200区间测试)
threshold = 180
image = image.point(lambda x: 0 if x < threshold else 255, '1')

针对列格式设置字符白名单

利用Tesseract的tessedit_char_whitelist参数,约束各列的识别字符范围,避免无关字符干扰:

  • 实际数量列(格式##,###):仅允许数字和逗号
  • 允许湿度列(格式##%):仅允许数字和百分号

如果能定位到表格的列坐标,可截取列区域单独识别,精度会更高:

# 示例:截取实际数量列区域(需根据表格实际坐标调整x1,y1,x2,y2)
qty_region = image.crop((x1, y1, x2, y2))
# 配置数量列专属识别规则
qty_config = r'--psm 7 --oem 3 -c tessedit_char_whitelist=0123456789,'
qty_text = pytesseract.image_to_string(qty_region, config=qty_config)

# 湿度列同理
humidity_region = image.crop((x3, y3, x4, y4))
humidity_config = r'--psm 7 --oem 3 -c tessedit_char_whitelist=0123456789%'
humidity_text = pytesseract.image_to_string(humidity_region, config=humidity_config)

调整PSM模式

针对单行/单列文本,使用--psm 7(将图像视为单行文本)替代--psm 6,减少上下文无关内容的干扰。


2. 利用已知集装箱号列表优化识别

步骤1:提取疑似集装箱号候选

先通过正则匹配,从OCR结果中筛选符合客户集装箱号格式的字符串(示例为4大写字母+7数字,需根据实际格式调整):

import re
# 自定义集装箱号格式正则
container_pattern = re.compile(r'[A-Z]{4}\d{7}')
for line in lines:
    candidates = container_pattern.findall(line)
    if candidates:
        raw_container = candidates[0]

步骤2:模糊匹配已知列表

使用编辑距离算法,在已知列表中找到最相似的匹配,推荐用fuzzywuzzy库实现:

from fuzzywuzzy import process

# 替换为客户实际的集装箱号列表
known_containers = ['ABCU1234567', 'DEFV9876543', ...]

# 匹配相似度最高的结果,设置阈值过滤低匹配度结果
best_match, score = process.extractOne(raw_container, known_containers)
if score >= 80:
    corrected_container = best_match
else:
    # 匹配度过低标记为待人工审核
    corrected_container = f"待审核:{raw_container}"

安装依赖:pip install fuzzywuzzy python-Levenshtein


3. 批量处理数百张图片

封装单图处理函数

将单张图片的预处理、OCR识别、数据提取逻辑封装为可复用函数:

import pandas as pd
import os

def process_single_image(image_path, known_containers):
    # 图像预处理
    image = Image.open(image_path)
    enhancer = ImageEnhance.Contrast(image)
    image = enhancer.enhance(1.5)
    image = image.filter(ImageFilter.MedianFilter())
    image = image.convert('L')
    threshold = 180
    image = image.point(lambda x: 0 if x < threshold else 255, '1')
    
    # OCR识别
    config = r'--psm 6 --oem 3 preserve_interword_spaces=1'
    extracted_text = pytesseract.image_to_string(image, lang='eng', config=config)
    lines = extracted_text.strip().split('\n')
    
    # 提取并整理数据(需根据表格实际列顺序调整)
    row_data = []
    for line in lines:
        if not line.strip():
            continue
        cols = line.split()
        # 处理集装箱号
        raw_container = cols[0]
        best_match, score = process.extractOne(raw_container, known_containers)
        corrected_container = best_match if score >=80 else f"待审核:{raw_container}"
        # 提取其他列
        row_data.append({
            "集装箱号": corrected_container,
            "实际数量": cols[1],
            "允许湿度": cols[3],
            "来源图片": os.path.basename(image_path)
        })
    return row_data

遍历文件夹批量处理

# 设置存放扫描件的文件夹路径
image_folder = "索赔报告扫描件"
# 获取所有图片文件(支持png、jpg、jpeg格式)
image_paths = [
    os.path.join(image_folder, filename) 
    for filename in os.listdir(image_folder) 
    if filename.lower().endswith(('.png', '.jpg', '.jpeg'))
]

# 批量处理所有图片
all_report_data = []
known_containers = ['ABCU1234567', 'DEFV9876543', ...]  # 替换为实际列表
for path in image_paths:
    print(f"正在处理:{path}")
    data = process_single_image(path, known_containers)
    all_report_data.extend(data)

# 将结果导出到Excel
df = pd.DataFrame(all_report_data)
df.to_excel("索赔报告汇总.xlsx", index=False)

内容的提问来源于stack exchange,提问作者Lefteris Kyprianou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 01:57:43