如何解决Python训练Hugging Face NER模型时的维度不匹配ValueError
解决NER训练中双GPU下的维度不匹配问题
问题重现
训练基于BERT+CRF的自定义NER模型,输入已完成截断和padding处理,使用2块GPU训练时触发如下维度不匹配错误:
ValueError: Caught ValueError in replica 0 on device 0. Original Traceback (most recent call last): File "/opt/conda/lib/python3.7/site-packages/torch/nn/parallel/parallel_apply.py", line 61, in _worker output = module(*input, **kwargs) File "/opt/conda/lib/python3.7/site-packages/torch/nn/modules/module.py", line 727, in _call_impl result = self.forward(*input, **kwargs) File "/tmp/ipykernel_1511906/3825976416.py", line 253, in forward return [loss.to(device), torch.tensor(prediction).to(device)] ValueError: expected sequence of length 15 at dim 1 (got 18)
问题原因
- CRF输出序列长度不统一:
crf.decode()返回的是每个样本的有效token标签序列(长度等于该样本attention_mask中1的数量),而非padding后的统一长度,直接转为张量会因维度不一致报错。 - 多GPU训练的硬性要求:分布式训练时,每个GPU返回的输出必须是形状完全统一的张量,不能是变长列表。
- 代码笔误:返回语句中使用了未定义的变量
prediction2,实际应为prediction。
修复方案
1. 对CRF预测结果做统一padding
将crf.decode()返回的变长序列填充到当前batch的最大长度(即attention_mask的第二维度长度),填充值可选用label2id['O'](无效标签ID)。
2. 修正变量拼写错误
将prediction2改为实际定义的变量prediction。
3. 可选优化:训练阶段仅返回loss
如果训练时不需要输出预测结果,可简化forward方法,只返回loss,减少不必要的张量操作。
修改后的关键代码
完整修正的BERT_CRF forward方法
def forward(self, input_ids, attention_mask, labels=None, token_type_ids=None): print("the types in forward",type(input_ids), type(attention_mask), type(labels),type(token_type_ids)) outputs = self.bert(input_ids, attention_mask=attention_mask) sequence_output = torch.stack((outputs[1][-1], outputs[1][-2], outputs[1][-3], outputs[1][-4])).mean(dim=0) sequence_output = self.dropout(sequence_output) emission = self.classifier(sequence_output) labels = labels.reshape(attention_mask.size()[0], attention_mask.size()[1]) if labels is not None: loss = -self.crf(log_soft(emission, 2), labels, mask=attention_mask.type(torch.uint8), reduction='mean') # 对CRF预测结果做统一padding predictions = self.crf.decode(emission, mask=attention_mask.type(torch.uint8)) max_len = attention_mask.size(1) padded_predictions = [pred + [label2id['O']]*(max_len - len(pred)) for pred in predictions] padded_predictions = torch.tensor(padded_predictions).to(device) # 修正变量拼写错误 return [loss.to(device), padded_predictions] else: predictions = self.crf.decode(emission, mask=attention_mask.type(torch.uint8)) max_len = attention_mask.size(1) padded_predictions = [pred + [label2id['O']]*(max_len - len(pred)) for pred in predictions] return torch.tensor(padded_predictions).to(device)
简化版(训练仅返回loss)
def forward(self, input_ids, attention_mask, labels=None, token_type_ids=None): outputs = self.bert(input_ids, attention_mask=attention_mask) sequence_output = torch.stack((outputs[1][-1], outputs[1][-2], outputs[1][-3], outputs[1][-4])).mean(dim=0) sequence_output = self.dropout(sequence_output) emission = self.classifier(sequence_output) labels = labels.reshape(attention_mask.size()[0], attention_mask.size()[1]) if labels is not None: # 训练阶段仅返回loss,满足Trainer反向传播需求 return -self.crf(log_soft(emission, 2), labels, mask=attention_mask.type(torch.uint8), reduction='mean') else: predictions = self.crf.decode(emission, mask=attention_mask.type(torch.uint8)) max_len = attention_mask.size(1) padded_predictions = [pred + [label2id['O']]*(max_len - len(pred)) for pred in predictions] return torch.tensor(padded_predictions).to(device)
额外注意事项
- 必须完成
tokenize_and_align_labels函数的实现,确保labels与input_ids长度一致,且标签与token正确对齐(当前代码中该函数为pass状态,无法正常生成训练数据)。 - 多GPU训练时,所有输入、输出张量的形状必须严格统一,禁止出现变长序列。
内容的提问来源于stack exchange,提问作者MAC
相关产品推荐
相关产品推荐

