基于Python Tesseract优化扫描单据表格数据提取的技术问询
针对扫描索赔报告OCR处理的优化方案
1. 针对固定格式列优化识别精度
强化图像预处理
在现有预处理流程基础上,增加二值化处理,强化文字与背景的边界,减少识别干扰:
# 转灰度后添加二值化步骤 image = image.convert('L') # 自定义阈值(可根据图片清晰度调整,建议160-200区间测试) threshold = 180 image = image.point(lambda x: 0 if x < threshold else 255, '1')
针对列格式设置字符白名单
利用Tesseract的tessedit_char_whitelist参数,约束各列的识别字符范围,避免无关字符干扰:
- 实际数量列(格式
##,###):仅允许数字和逗号 - 允许湿度列(格式
##%):仅允许数字和百分号
如果能定位到表格的列坐标,可截取列区域单独识别,精度会更高:
# 示例:截取实际数量列区域(需根据表格实际坐标调整x1,y1,x2,y2) qty_region = image.crop((x1, y1, x2, y2)) # 配置数量列专属识别规则 qty_config = r'--psm 7 --oem 3 -c tessedit_char_whitelist=0123456789,' qty_text = pytesseract.image_to_string(qty_region, config=qty_config) # 湿度列同理 humidity_region = image.crop((x3, y3, x4, y4)) humidity_config = r'--psm 7 --oem 3 -c tessedit_char_whitelist=0123456789%' humidity_text = pytesseract.image_to_string(humidity_region, config=humidity_config)
调整PSM模式
针对单行/单列文本,使用--psm 7(将图像视为单行文本)替代--psm 6,减少上下文无关内容的干扰。
2. 利用已知集装箱号列表优化识别
步骤1:提取疑似集装箱号候选
先通过正则匹配,从OCR结果中筛选符合客户集装箱号格式的字符串(示例为4大写字母+7数字,需根据实际格式调整):
import re # 自定义集装箱号格式正则 container_pattern = re.compile(r'[A-Z]{4}\d{7}') for line in lines: candidates = container_pattern.findall(line) if candidates: raw_container = candidates[0]
步骤2:模糊匹配已知列表
使用编辑距离算法,在已知列表中找到最相似的匹配,推荐用fuzzywuzzy库实现:
from fuzzywuzzy import process # 替换为客户实际的集装箱号列表 known_containers = ['ABCU1234567', 'DEFV9876543', ...] # 匹配相似度最高的结果,设置阈值过滤低匹配度结果 best_match, score = process.extractOne(raw_container, known_containers) if score >= 80: corrected_container = best_match else: # 匹配度过低标记为待人工审核 corrected_container = f"待审核:{raw_container}"
安装依赖:pip install fuzzywuzzy python-Levenshtein
3. 批量处理数百张图片
封装单图处理函数
将单张图片的预处理、OCR识别、数据提取逻辑封装为可复用函数:
import pandas as pd import os def process_single_image(image_path, known_containers): # 图像预处理 image = Image.open(image_path) enhancer = ImageEnhance.Contrast(image) image = enhancer.enhance(1.5) image = image.filter(ImageFilter.MedianFilter()) image = image.convert('L') threshold = 180 image = image.point(lambda x: 0 if x < threshold else 255, '1') # OCR识别 config = r'--psm 6 --oem 3 preserve_interword_spaces=1' extracted_text = pytesseract.image_to_string(image, lang='eng', config=config) lines = extracted_text.strip().split('\n') # 提取并整理数据(需根据表格实际列顺序调整) row_data = [] for line in lines: if not line.strip(): continue cols = line.split() # 处理集装箱号 raw_container = cols[0] best_match, score = process.extractOne(raw_container, known_containers) corrected_container = best_match if score >=80 else f"待审核:{raw_container}" # 提取其他列 row_data.append({ "集装箱号": corrected_container, "实际数量": cols[1], "允许湿度": cols[3], "来源图片": os.path.basename(image_path) }) return row_data
遍历文件夹批量处理
# 设置存放扫描件的文件夹路径 image_folder = "索赔报告扫描件" # 获取所有图片文件(支持png、jpg、jpeg格式) image_paths = [ os.path.join(image_folder, filename) for filename in os.listdir(image_folder) if filename.lower().endswith(('.png', '.jpg', '.jpeg')) ] # 批量处理所有图片 all_report_data = [] known_containers = ['ABCU1234567', 'DEFV9876543', ...] # 替换为实际列表 for path in image_paths: print(f"正在处理:{path}") data = process_single_image(path, known_containers) all_report_data.extend(data) # 将结果导出到Excel df = pd.DataFrame(all_report_data) df.to_excel("索赔报告汇总.xlsx", index=False)
内容的提问来源于stack exchange,提问作者Lefteris Kyprianou
相关产品推荐
相关产品推荐

