使用LayoutLMv3Processor传递is_split_into_words参数时触发类型错误的问题求助
LayoutLMv3Processor传递is_split_into_words参数时触发类型错误的问题求助
大家好,我目前正在用HuggingFace Transformers库微调LayoutLMv3模型做token classification任务。为了确保标签和分词后的token能正确对齐,我想在预处理阶段使用is_split_into_words=True这个参数,但调用processor时遇到了报错,希望能得到大家的帮助。
我的基本配置是:使用LayoutLMv3Processor,数据包含单词列表(对应example["words"])、边界框(bboxes)以及标签。尝试在调用processor时传入is_split_into_words=True,但触发了类型错误。
以下是我的预处理代码:
from transformers import LayoutLMv3Processor, LayoutLMv3Tokenizer from PIL import Image # 初始化分词器和processor tokenizer = LayoutLMv3Tokenizer.from_pretrained("microsoft/layoutlmv3-base") processor = LayoutLMv3Processor.from_pretrained("microsoft/layoutlmv3-base") def normalize_bbox(bbox, width, height): # 自定义边界框归一化逻辑 return [ int(1000 * (bbox[0] / width)), int(1000 * (bbox[1] / height)), int(1000 * (bbox[2] / width)), int(1000 * (bbox[3] / height)), ] def preprocess(example): image = Image.open(example["image_path"]).convert("RGB") image_width, image_height = image.size normalized_bboxes = [normalize_bbox(bbox, image_width, image_height) for bbox in example["bboxes"]] encoding = processor( image, example["words"], is_split_into_words=True, boxes=normalized_bboxes, word_labels=[label2id[l] for l in example["labels"]], truncation=True, padding="max_length", return_tensors="pt" ) return { "input_ids": encoding["input_ids"].squeeze(0), "attention_mask": encoding["attention_mask"].squeeze(0), "bbox": encoding["bbox"].squeeze(0), "pixel_values": encoding["pixel_values"].squeeze(0), "labels": encoding["labels"].squeeze(0) } # 处理数据集 tokenized_dataset = dataset.map(preprocess, remove_columns=dataset.column_names)
运行后触发的错误信息:
TypeError: LayoutLMv3TokenizerFast._batch_encode_plus() got an unexpected keyword argument 'is_split_into_words'
我排查后发现,似乎是LayoutLMv3Processor.__call__()方法没有把is_split_into_words参数转发给内部的分词器,导致分词器无法识别这个参数。有没有朋友遇到过类似问题?或者知道该如何正确传递这个参数来实现标签对齐呢?
内容来源于stack exchange
相关产品推荐
相关产品推荐

