You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解析DocTr/easyOCR的OCR结果,实现字段值映射与CSV构建?

解决W2表单解析的两个问题

一、将DocTr的孤立单词转为「字段名: 值」结构化输出

DocTr返回的孤立单词需要结合W2表单的固定布局规则关联标签和值,核心是利用文本的坐标位置分组匹配:

  1. 按行聚类文本
    W2表单的字段标签和对应值基本在同一水平行,可根据单词的y轴坐标(bbox的y1/y2)将同一行的单词归为一组,设置一个阈值(比如10像素)判断两个单词是否属于同一行。

  2. 区分标签与值

  • 同一行内,左侧文本通常是字段标签(如Employee's name、SSN),右侧为对应值;
  • 提前定义W2标准字段列表(比如["Employee's name", "Social security number", "Wages, tips, other compensation", ...]),匹配每行中的标签文本,剩余部分即为对应值。
  1. 代码实现示例
from doctr.models import ocr_predictor

# 假设已获取DocTr的Document对象
doc = ocr_predictor("path/to/w2_sample.jpg")
page_words = doc.pages[0].words

# 按y坐标聚类行
rows = {}
y_threshold = 10
for word in page_words:
    text = word.text
    y_center = (word.bbox[1] + word.bbox[3]) / 2
    # 找到对应的行组
    matched_row = None
    for row_y in rows:
        if abs(y_center - row_y) < y_threshold:
            matched_row = row_y
            break
    if matched_row:
        rows[matched_row].append((word.bbox[0], text))  # 存x坐标和文本,方便排序
    else:
        rows[y_center] = [(word.bbox[0], text)]

# 处理每行,生成结构化字典
w2_data = {}
# 预定义W2核心字段
w2_fields = ["Employee's name", "Social security number", "Wages, tips, other compensation", "Federal income tax withheld", "State wages, tips, etc.", "State income tax withheld"]

for row in rows.values():
    # 按x坐标排序,确保从左到右的顺序
    row_sorted = sorted(row, key=lambda x: x[0])
    row_text = " ".join([t for _, t in row_sorted])
    # 匹配字段标签
    for field in w2_fields:
        if field.lower() in row_text.lower():
            # 提取值:去掉标签部分,保留剩余内容
            value = row_text.replace(field, "").strip()
            w2_data[field] = value
            break

# 输出结构化结果
for field, value in w2_data.items():
    print(f"{field}: {value}")

二、解析easyOCR输出构建CSV

easyOCR的输出格式为[[(bbox), text, confidence], ...],同样利用坐标分组结构化数据并生成CSV:

  1. 按行整理文本
    和DocTr的处理逻辑一致,通过y轴坐标聚类同一行的文本,按x轴排序保证顺序。

  2. 映射到CSV列
    提前定义CSV表头(对应W2标准字段),将每行匹配到的标签和值对应到表头列,最后写入CSV文件。

  3. 代码实现示例

import easyocr
import csv

# 运行easyOCR获取结果
reader = easyocr.Reader(['en'])
result = reader.readtext("path/to/w2_sample.jpg")

# 按行聚类
rows = {}
y_threshold = 10
for (bbox, text, conf) in result:
    y_center = (bbox[0][1] + bbox[2][1]) / 2
    matched_row = None
    for row_y in rows:
        if abs(y_center - row_y) < y_threshold:
            matched_row = row_y
            break
    if matched_row:
        rows[matched_row].append((bbox[0][0], text))
    else:
        rows[y_center] = [(bbox[0][0], text)]

# 构建CSV数据
csv_header = ["Employee Name", "SSN", "Federal Wages", "Federal Tax Withheld", "State Wages", "State Tax Withheld"]
csv_row = [""] * len(csv_header)

field_mapping = {
    "Employee's name": 0,
    "Social security number": 1,
    "Wages, tips, other compensation": 2,
    "Federal income tax withheld": 3,
    "State wages, tips, etc.": 4,
    "State income tax withheld": 5
}

for row in rows.values():
    row_sorted = sorted(row, key=lambda x: x[0])
    row_text = " ".join([t for _, t in row_sorted])
    for field, idx in field_mapping.items():
        if field.lower() in row_text.lower():
            value = row_text.replace(field, "").strip()
            csv_row[idx] = value
            break

# 写入CSV文件
with open("w2_output.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.writer(f)
    writer.writerow(csv_header)
    writer.writerow(csv_row)

注意事项

  • 若表单存在多行值(比如员工地址),需调整聚类逻辑,允许同一字段跨多行匹配;
  • 对于模糊或识别错误的文本,可增加置信度过滤(比如只保留confidence>0.7的结果);
  • 不同版本的W2布局可能略有差异,可根据实际样本调整字段匹配规则和坐标阈值。

内容的提问来源于stack exchange,提问作者Raj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 15:27:18