You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于CORD-v2数据集微调LayoutLMv3的TTC提取预处理问题

解决CORD-v2数据集预处理中提取总金额(TTC)的问题

核心思路

CORD-v2已直接提供文本实体的真值标注,无需从图像重新OCR提取总金额,预处理的核心是将标注数据适配LayoutLMv3的输入格式,而非从零提取实体。


步骤1:解析CORD-v2的真值标注结构

CORD-v2每个图像对应同名JSON标注文件,核心字段为text_blocks数组,每个元素包含:

  • text:文本块内容
  • bbox:文本块的边界框坐标(格式为[left, top, right, bottom])
  • label:实体类型,总金额对应的标签为total

直接遍历筛选即可定位总金额:

import json

def extract_total_amount(ann_path):
    with open(ann_path, "r", encoding="utf-8") as f:
        annotation = json.load(f)
    # 筛选总金额文本块
    total_blocks = [block for block in annotation["text_blocks"] if block["label"] == "total"]
    # 返回文本和边界框(若有多个取第一个,根据数据集实际情况调整)
    return total_blocks[0]["text"] if total_blocks else None, total_blocks[0]["bbox"] if total_blocks else None

步骤2:适配LayoutLMv3的输入格式

LayoutLMv3需要图像像素、文本token、token级bbox、任务标签四个核心输入,处理流程如下:

  1. 图像预处理:用模型配套的LayoutLMv3ImageProcessor做resize、归一化:
    from transformers import LayoutLMv3ImageProcessor
    from PIL import Image
    
    image_processor = LayoutLMv3ImageProcessor(apply_ocr=False)  # 无需OCR,用真值文本
    img = Image.open("sample_image.jpg")
    pixel_values = image_processor(img, return_tensors="pt").pixel_values
    
  2. 文本与bbox映射:将所有文本块按空间顺序拼接为文本序列,同时为每个token分配对应文本块的bbox(需归一化到0-1000范围):
    from transformers import LayoutLMv3Tokenizer
    
    tokenizer = LayoutLMv3Tokenizer.from_pretrained("microsoft/layoutlmv3-base")
    
    def prepare_text_and_bboxes(annotation, img):
        texts = []
        bboxes = []
        # 按y坐标(top)排序,保证空间顺序
        sorted_blocks = sorted(annotation["text_blocks"], key=lambda x: x["bbox"][1])
        for block in sorted_blocks:
            texts.append(block["text"])
            # 归一化bbox到0-1000(LayoutLMv3要求)
            norm_bbox = [
                int(1000 * coord / img.size[0]) if i % 2 == 0 else int(1000 * coord / img.size[1])
                for i, coord in enumerate(block["bbox"])
            ]
            bboxes.append(norm_bbox)
        # 拼接文本并token化,同时映射bbox到每个token
        encoding = tokenizer(texts, padding="max_length", truncation=True, max_length=512, return_offsets_mapping=True)
        # 为每个token分配对应文本块的bbox
        token_bboxes = []
        for i, offset in enumerate(encoding["offset_mapping"]):
            if offset[0] == 0 and offset[1] == 0:
                token_bboxes.append([0,0,0,0])  # padding token的bbox
            else:
                # 找到当前token所属的文本块
                block_idx = next(j for j, text in enumerate(texts) if offset[0] <= len(text))
                token_bboxes.append(bboxes[block_idx])
        encoding["bbox"] = token_bboxes
        return encoding
    
  3. 任务标签构建:用IOB标注格式标记总金额对应的token,其他token标记为O:
    def prepare_labels(encoding, annotation):
        total_text = extract_total_amount(annotation)[0]
        labels = ["O"] * len(encoding["input_ids"])
        if not total_text:
            return labels
        # 找到总金额文本在token序列中的位置
        total_tokens = tokenizer.encode(total_text, add_special_tokens=False)
        # 匹配token序列中的总金额片段
        for i in range(len(encoding["input_ids"]) - len(total_tokens)):
            if encoding["input_ids"][i:i+len(total_tokens)] == total_tokens:
                labels[i] = "B-TOTAL"
                for j in range(1, len(total_tokens)):
                    labels[i+j] = "I-TOTAL"
                break
        return labels
    

步骤3:解决旧版CORD示例的兼容问题

旧版CORD的标注字段与v2差异较大,需修改代码中的核心映射:

  • 把旧代码中调用regions的部分替换为text_blocks
  • 把旧代码中实体标签(如total_amount)替换为CORD-v2的total
  • 关闭OCR逻辑(旧版可能依赖OCR提取文本,v2直接用真值即可)

调试技巧

  1. 打印标注结构确认字段:
    with open("sample_annotation.json", "r") as f:
        ann = json.load(f)
    print([(b["text"], b["label"]) for b in ann["text_blocks"]][:5])  # 查看前5个文本块
    
  2. 可视化总金额的bbox:
    from PIL import Image, ImageDraw
    img = Image.open("sample_image.jpg")
    draw = ImageDraw.Draw(img)
    total_bbox = extract_total_amount("sample_annotation.json")[1]
    draw.rectangle(total_bbox, outline="red", width=2)
    img.show()
    
  3. 检查模型输入的格式合法性:确认bbox数值在0-1000之间,token数量不超过模型最大长度(默认512)

内容的提问来源于stack exchange,提问作者DRISSI EL Houcine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 10:29:54