LayoutLMv2推理时张量尺寸不匹配RuntimeError问题求助
问题描述
基于HuggingFace训练LayoutLMv2模型后,单图推理时报错:
RuntimeError: The expanded size of the tensor (1011) must match the existing size (512) at non-singleton dimension 1. Target sizes: [1, 1011]. Tensor sizes: [1, 512]
错误发生在model(**encoded_inputs)执行阶段。
推理代码
query = '/Users/vaihabsaxena/Desktop/Newfolder/labeled/Others/Two.pdf26.png' image = Image.open(query).convert("RGB") encoded_inputs = processor(image, return_tensors="pt").to(device) outputs = model(**encoded_inputs) preds = torch.softmax(outputs.logits, dim=1).tolist()[0] pred_labels = {label:pred for label, pred in zip(label2idx.keys(), preds)} pred_labels
Processor初始化代码
feature_extractor = LayoutLMv2FeatureExtractor() tokenizer = LayoutLMv2Tokenizer.from_pretrained("microsoft/layoutlmv2-base-uncased") processor = LayoutLMv2Processor(feature_extractor, tokenizer)
模型训练核心代码
model = LayoutLMv2ForSequenceClassification.from_pretrained( "microsoft/layoutlmv2-base-uncased", num_labels=len(label2idx) ) model.to(device); # 训练循环省略...
完整错误堆栈
RuntimeError Traceback (most recent call last) /Users/vaihabsaxena/Desktop/Newfolder/pytorch.ipynb Cell 37 in <cell line: 4>() 2 image = Image.open(query).convert("RGB") 3 encoded_inputs = processor(image, return_tensors="pt").to(device) ----> 4 outputs = model(**encoded_inputs) 5 preds = torch.softmax(outputs.logits, dim=1).tolist()[0] 6 pred_labels = {label:pred for label, pred in zip(label2idx.keys(), preds)} File ~/opt/anaconda3/envs/env_pytorch/lib/python3.9/site-packages/torch/nn/modules/module.py:1130, in Module._call_impl(self, *input, **kwargs) 1126 # If we don't have any hooks, we want to skip the rest of the logic in 1127 # this function, and just call forward. 1128 if not (self._backward_hooks or self._forward_hooks or self._forward_pre_hooks or _global_backward_hooks 1129 or _global_forward_hooks or _global_forward_pre_hooks): -> 1130 return forward_call(*input, **kwargs) 1131 # Do not call functions when jit is used 1132 full_backward_hooks, non_full_backward_hooks = [], [] File ~/opt/anaconda3/envs/env_pytorch/lib/python3.9/site-packages/transformers/models/layoutlmv2/modeling_layoutlmv2.py:1071, in LayoutLMv2ForSequenceClassification.forward(self, input_ids, bbox, image, attention_mask, token_type_ids, position_ids, head_mask, inputs_embeds, labels, output_attentions, output_hidden_states, return_dict) 1061 visual_position_ids = torch.arange(0, visual_shape[1], dtype=torch.long, device=device).repeat( 1062 input_shape[0], 1 1063 ) 1065 initial_image_embeddings = self.layoutlmv2._calc_img_embeddings( 1066 image=image, 1067 bbox=visual_bbox, ... 896 input_shape[0], 1 897 ) 898 final_position_ids = torch.cat([position_ids, visual_position_ids], dim=1) RuntimeError: The expanded size of the tensor (1011) must match the existing size (512) at non-singleton dimension 1. Target sizes: [1, 1011]. Tensor sizes: [1, 512]
尝试单独设置tokenizer截断最大长度时,出现encoded_inputs为None的问题,求排查原因与解决方法。
问题原因
- 输入序列长度不匹配:
microsoft/layoutlmv2-base-uncased默认最大序列长度为512,推理时processor处理图片生成的token序列长度(1011)超过模型输入限制,而训练阶段的dataloader自动做了截断/填充,导致训练与推理的输入维度不一致,触发张量拼接错误。 - 截断配置逻辑错误:单独修改tokenizer的截断参数无效,LayoutLMv2的输入处理由
LayoutLMv2Processor统一协调文本token与视觉特征,单独修改tokenizer不会同步到processor的处理流程,导致输入处理异常返回None。
解决方法
1. 初始化Processor时统一配置截断与填充
修改processor初始化代码,明确指定最大长度、截断和填充策略,确保训练与推理的输入处理逻辑一致:
feature_extractor = LayoutLMv2FeatureExtractor() tokenizer = LayoutLMv2Tokenizer.from_pretrained("microsoft/layoutlmv2-base-uncased") # 统一设置最大长度、截断、填充规则 processor = LayoutLMv2Processor( feature_extractor, tokenizer, max_length=512, truncation=True, padding="max_length" )
2. 推理阶段正确使用Processor
推理时直接调用processor即可,无需重复设置参数(已在初始化时统一配置):
query = '/Users/vaihabsaxena/Desktop/Newfolder/labeled/Others/Two.pdf26.png' image = Image.open(query).convert("RGB") # processor自动将输入截断/填充到512长度 encoded_inputs = processor(image, return_tensors="pt").to(device) # 可选:验证输入维度是否符合要求 print("input_ids shape:", encoded_inputs["input_ids"].shape) # 应为 [1, 512] outputs = model(**encoded_inputs) preds = torch.softmax(outputs.logits, dim=1).tolist()[0] pred_labels = {label:pred for label, pred in zip(label2idx.keys(), preds)} pred_labels
3. 对齐训练与推理的输入配置
如果训练时的dataloader使用了自定义截断/填充参数,需确保processor的初始化参数与dataloader配置完全一致,避免训练与推理的输入分布差异。
内容的提问来源于stack exchange,提问作者Vai
相关产品推荐
相关产品推荐

