You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何制作适配HuggingFace Transformers与Trainer的多头回归DataLoader

多头回归任务Trainer报错ValueError的解决

问题背景

做多头回归任务,每个文本要预测5个分数。已配置problem_type = 'regression'、模型num_labels=5,但用Trainer训练时触发ValueError,提示标签过度嵌套。之前单标签回归(num_classes=1)时正常运行。

错误日志

raise ValueError(
ValueError: Unable to create tensor, you should probably activate truncation and/or padding with 'padding=True' 'truncation=True' to have batched tensors with the same length. Perhaps your features (`labels` in this case) have excessive nesting (inputs type `list` where type `int` is expected).

模型代码

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
model = AutoModelForSequenceClassification.from_pretrained("bert-base-uncased", 
                                                           num_labels=5, 
                                                          problem_type = "regression")

自定义Dataset代码

class MultiRegressionDataset(torch.utils.data.Dataset):
    def __init__(self, texts, labels):
        self.labels = labels
        self.texts = texts
        

    def __getitem__(self, idx, sanity_check = False):
        output = tokenizer(self.texts[idx], truncation=True,
                              padding="max_length",
                              max_length = 128) # 返回字典

        output['labels'] = torch.tensor(self.labels[idx])
        
        return output

data = MultiRegressionDataset(["text1", "text2"], [[1,2,3,4,5], [5,4,3,2,1]])

data.__getitem__(0) # 可正常获取单样本数据

已尝试的无效方案

  • output['labels'] = torch.tensor(self.labels[idx]).unsqueeze(-1)
  • 结合return_tensors = "pt"执行上述操作

问题原因与解决方法

核心问题

  1. 标签格式不统一:tokenizer返回的input_ids等是Python列表,而你提前把标签转成了张量,导致DataLoader在拼批时无法兼容两种不同类型的数据,引发嵌套错误。
  2. 错误的维度调整:unsqueeze(-1)把原本符合要求的(5,)一维标签变成了(5,1)二维张量,反而增加了嵌套层次,不符合模型对多头回归标签的维度要求。

修正后的Dataset代码

class MultiRegressionDataset(torch.utils.data.Dataset):
    def __init__(self, texts, labels):
        self.labels = labels
        self.texts = texts
        

    def __getitem__(self, idx, sanity_check = False):
        output = tokenizer(self.texts[idx], truncation=True,
                              padding="max_length",
                              max_length = 128)
        # 关键:直接返回标签列表,不要提前转张量
        output['labels'] = self.labels[idx]
        
        return output

关键说明

  • 不用在Dataset中提前转换标签为张量,Trainer会自动将批量的二维列表标签转换为形状为(batch_size, 5)的张量,完全匹配num_labels=5的回归模型输入要求。
  • 确保你的标签数据本身就是二维结构:每个样本对应一个长度为5的列表,比如示例中的[[1,2,3,4,5], [5,4,3,2,1]],这是正确的输入格式。

内容的提问来源于stack exchange,提问作者Deshwal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 15:33:44