You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Python训练Hugging Face NER模型时的维度不匹配ValueError

解决NER训练中双GPU下的维度不匹配问题

问题重现

训练基于BERT+CRF的自定义NER模型,输入已完成截断和padding处理,使用2块GPU训练时触发如下维度不匹配错误:

ValueError: Caught ValueError in replica 0 on device 0.
Original Traceback (most recent call last):
  File "/opt/conda/lib/python3.7/site-packages/torch/nn/parallel/parallel_apply.py", line 61, in _worker
    output = module(*input, **kwargs)
  File "/opt/conda/lib/python3.7/site-packages/torch/nn/modules/module.py", line 727, in _call_impl
    result = self.forward(*input, **kwargs)
  File "/tmp/ipykernel_1511906/3825976416.py", line 253, in forward
    return [loss.to(device), torch.tensor(prediction).to(device)]
ValueError: expected sequence of length 15 at dim 1 (got 18)

问题原因

  1. CRF输出序列长度不统一:crf.decode()返回的是每个样本的有效token标签序列(长度等于该样本attention_mask中1的数量),而非padding后的统一长度,直接转为张量会因维度不一致报错。
  2. 多GPU训练的硬性要求:分布式训练时,每个GPU返回的输出必须是形状完全统一的张量,不能是变长列表。
  3. 代码笔误:返回语句中使用了未定义的变量prediction2,实际应为prediction。

修复方案

1. 对CRF预测结果做统一padding

将crf.decode()返回的变长序列填充到当前batch的最大长度(即attention_mask的第二维度长度),填充值可选用label2id['O'](无效标签ID)。

2. 修正变量拼写错误

将prediction2改为实际定义的变量prediction。

3. 可选优化:训练阶段仅返回loss

如果训练时不需要输出预测结果,可简化forward方法,只返回loss,减少不必要的张量操作。

修改后的关键代码

完整修正的BERT_CRF forward方法

def forward(self, input_ids, attention_mask, labels=None, token_type_ids=None):
    print("the types in forward",type(input_ids), type(attention_mask), type(labels),type(token_type_ids))
    
    outputs = self.bert(input_ids, attention_mask=attention_mask)

    sequence_output = torch.stack((outputs[1][-1], outputs[1][-2], outputs[1][-3], outputs[1][-4])).mean(dim=0)
    sequence_output = self.dropout(sequence_output)
    emission = self.classifier(sequence_output)  

    labels = labels.reshape(attention_mask.size()[0], attention_mask.size()[1])

    if labels is not None:
        loss = -self.crf(log_soft(emission, 2), labels, mask=attention_mask.type(torch.uint8), reduction='mean')
        
        # 对CRF预测结果做统一padding
        predictions = self.crf.decode(emission, mask=attention_mask.type(torch.uint8))
        max_len = attention_mask.size(1)
        padded_predictions = [pred + [label2id['O']]*(max_len - len(pred)) for pred in predictions]
        padded_predictions = torch.tensor(padded_predictions).to(device)
        
        # 修正变量拼写错误
        return [loss.to(device), padded_predictions]
    else:
        predictions = self.crf.decode(emission, mask=attention_mask.type(torch.uint8))
        max_len = attention_mask.size(1)
        padded_predictions = [pred + [label2id['O']]*(max_len - len(pred)) for pred in predictions]
        return torch.tensor(padded_predictions).to(device)

简化版(训练仅返回loss)

def forward(self, input_ids, attention_mask, labels=None, token_type_ids=None):
    outputs = self.bert(input_ids, attention_mask=attention_mask)

    sequence_output = torch.stack((outputs[1][-1], outputs[1][-2], outputs[1][-3], outputs[1][-4])).mean(dim=0)
    sequence_output = self.dropout(sequence_output)
    emission = self.classifier(sequence_output)  

    labels = labels.reshape(attention_mask.size()[0], attention_mask.size()[1])

    if labels is not None:
        # 训练阶段仅返回loss,满足Trainer反向传播需求
        return -self.crf(log_soft(emission, 2), labels, mask=attention_mask.type(torch.uint8), reduction='mean')
    else:
        predictions = self.crf.decode(emission, mask=attention_mask.type(torch.uint8))
        max_len = attention_mask.size(1)
        padded_predictions = [pred + [label2id['O']]*(max_len - len(pred)) for pred in predictions]
        return torch.tensor(padded_predictions).to(device)

额外注意事项

  • 必须完成tokenize_and_align_labels函数的实现,确保labels与input_ids长度一致,且标签与token正确对齐(当前代码中该函数为pass状态,无法正常生成训练数据)。
  • 多GPU训练时,所有输入、输出张量的形状必须严格统一,禁止出现变长序列。

内容的提问来源于stack exchange,提问作者MAC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 16:23:08