You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用LayoutLMv3Processor传递is_split_into_words参数时触发类型错误的问题求助

LayoutLMv3Processor传递is_split_into_words参数时触发类型错误的问题求助

大家好,我目前正在用HuggingFace Transformers库微调LayoutLMv3模型做token classification任务。为了确保标签和分词后的token能正确对齐,我想在预处理阶段使用is_split_into_words=True这个参数,但调用processor时遇到了报错,希望能得到大家的帮助。

我的基本配置是:使用LayoutLMv3Processor,数据包含单词列表(对应example["words"])、边界框(bboxes)以及标签。尝试在调用processor时传入is_split_into_words=True,但触发了类型错误。

以下是我的预处理代码:

from transformers import LayoutLMv3Processor, LayoutLMv3Tokenizer
from PIL import Image

# 初始化分词器和processor
tokenizer = LayoutLMv3Tokenizer.from_pretrained("microsoft/layoutlmv3-base")
processor = LayoutLMv3Processor.from_pretrained("microsoft/layoutlmv3-base")

def normalize_bbox(bbox, width, height):
    # 自定义边界框归一化逻辑
    return [
        int(1000 * (bbox[0] / width)),
        int(1000 * (bbox[1] / height)),
        int(1000 * (bbox[2] / width)),
        int(1000 * (bbox[3] / height)),
    ]

def preprocess(example):
    image = Image.open(example["image_path"]).convert("RGB")
    image_width, image_height = image.size
    normalized_bboxes = [normalize_bbox(bbox, image_width, image_height) for bbox in example["bboxes"]]
    
    encoding = processor(
        image,
        example["words"],
        is_split_into_words=True,
        boxes=normalized_bboxes,
        word_labels=[label2id[l] for l in example["labels"]],
        truncation=True,
        padding="max_length",
        return_tensors="pt"
    )
    
    return {
        "input_ids": encoding["input_ids"].squeeze(0),
        "attention_mask": encoding["attention_mask"].squeeze(0),
        "bbox": encoding["bbox"].squeeze(0),
        "pixel_values": encoding["pixel_values"].squeeze(0),
        "labels": encoding["labels"].squeeze(0)
    }

# 处理数据集
tokenized_dataset = dataset.map(preprocess, remove_columns=dataset.column_names)

运行后触发的错误信息:

TypeError: LayoutLMv3TokenizerFast._batch_encode_plus() got an unexpected keyword argument 'is_split_into_words'

我排查后发现,似乎是LayoutLMv3Processor.__call__()方法没有把is_split_into_words参数转发给内部的分词器,导致分词器无法识别这个参数。有没有朋友遇到过类似问题?或者知道该如何正确传递这个参数来实现标签对齐呢?

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 09:49:04