如何解析DocTr/easyOCR的OCR结果,实现字段值映射与CSV构建?
解决W2表单解析的两个问题
一、将DocTr的孤立单词转为「字段名: 值」结构化输出
DocTr返回的孤立单词需要结合W2表单的固定布局规则关联标签和值,核心是利用文本的坐标位置分组匹配:
按行聚类文本
W2表单的字段标签和对应值基本在同一水平行,可根据单词的y轴坐标(bbox的y1/y2)将同一行的单词归为一组,设置一个阈值(比如10像素)判断两个单词是否属于同一行。区分标签与值
- 同一行内,左侧文本通常是字段标签(如
Employee's name、SSN),右侧为对应值; - 提前定义W2标准字段列表(比如
["Employee's name", "Social security number", "Wages, tips, other compensation", ...]),匹配每行中的标签文本,剩余部分即为对应值。
- 代码实现示例
from doctr.models import ocr_predictor # 假设已获取DocTr的Document对象 doc = ocr_predictor("path/to/w2_sample.jpg") page_words = doc.pages[0].words # 按y坐标聚类行 rows = {} y_threshold = 10 for word in page_words: text = word.text y_center = (word.bbox[1] + word.bbox[3]) / 2 # 找到对应的行组 matched_row = None for row_y in rows: if abs(y_center - row_y) < y_threshold: matched_row = row_y break if matched_row: rows[matched_row].append((word.bbox[0], text)) # 存x坐标和文本,方便排序 else: rows[y_center] = [(word.bbox[0], text)] # 处理每行,生成结构化字典 w2_data = {} # 预定义W2核心字段 w2_fields = ["Employee's name", "Social security number", "Wages, tips, other compensation", "Federal income tax withheld", "State wages, tips, etc.", "State income tax withheld"] for row in rows.values(): # 按x坐标排序,确保从左到右的顺序 row_sorted = sorted(row, key=lambda x: x[0]) row_text = " ".join([t for _, t in row_sorted]) # 匹配字段标签 for field in w2_fields: if field.lower() in row_text.lower(): # 提取值:去掉标签部分,保留剩余内容 value = row_text.replace(field, "").strip() w2_data[field] = value break # 输出结构化结果 for field, value in w2_data.items(): print(f"{field}: {value}")
二、解析easyOCR输出构建CSV
easyOCR的输出格式为[[(bbox), text, confidence], ...],同样利用坐标分组结构化数据并生成CSV:
按行整理文本
和DocTr的处理逻辑一致,通过y轴坐标聚类同一行的文本,按x轴排序保证顺序。映射到CSV列
提前定义CSV表头(对应W2标准字段),将每行匹配到的标签和值对应到表头列,最后写入CSV文件。代码实现示例
import easyocr import csv # 运行easyOCR获取结果 reader = easyocr.Reader(['en']) result = reader.readtext("path/to/w2_sample.jpg") # 按行聚类 rows = {} y_threshold = 10 for (bbox, text, conf) in result: y_center = (bbox[0][1] + bbox[2][1]) / 2 matched_row = None for row_y in rows: if abs(y_center - row_y) < y_threshold: matched_row = row_y break if matched_row: rows[matched_row].append((bbox[0][0], text)) else: rows[y_center] = [(bbox[0][0], text)] # 构建CSV数据 csv_header = ["Employee Name", "SSN", "Federal Wages", "Federal Tax Withheld", "State Wages", "State Tax Withheld"] csv_row = [""] * len(csv_header) field_mapping = { "Employee's name": 0, "Social security number": 1, "Wages, tips, other compensation": 2, "Federal income tax withheld": 3, "State wages, tips, etc.": 4, "State income tax withheld": 5 } for row in rows.values(): row_sorted = sorted(row, key=lambda x: x[0]) row_text = " ".join([t for _, t in row_sorted]) for field, idx in field_mapping.items(): if field.lower() in row_text.lower(): value = row_text.replace(field, "").strip() csv_row[idx] = value break # 写入CSV文件 with open("w2_output.csv", "w", newline="", encoding="utf-8") as f: writer = csv.writer(f) writer.writerow(csv_header) writer.writerow(csv_row)
注意事项
- 若表单存在多行值(比如员工地址),需调整聚类逻辑,允许同一字段跨多行匹配;
- 对于模糊或识别错误的文本,可增加置信度过滤(比如只保留confidence>0.7的结果);
- 不同版本的W2布局可能略有差异,可根据实际样本调整字段匹配规则和坐标阈值。
内容的提问来源于stack exchange,提问作者Raj
相关产品推荐
相关产品推荐

