You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Transformers Trainer训练时模型未返回loss的问题求助

训练报错:模型未返回loss,标签丢失

完整错误信息

ValueError: The model did not return a loss from the inputs, only the following keys: logits. For reference, the inputs it received are input_ids,attention_mask.

问题背景

数据集本身包含label字段,传入Trainer的train_dataset的column_names也显示包含label(列表为['text', 'label', 'input_ids', 'attention_mask']),但训练时出现上述错误。

训练代码

from transformers import Trainer, TrainingArguments
batch_size = 64
logging_steps = len(emotions_encoded["train"])

training_args = TrainingArguments(output_dir = "model_out",
                                  num_train_epochs=2, 
                                  learning_rate = 2e-5, 
                                  per_device_train_batch_size=batch_size, 
                                  per_device_eval_batch_size=batch_size, 
                                  weight_decay=0.01, 
                                  evaluation_strategy="epoch", 
                                  disable_tqdm=False, 
                                  logging_steps=logging_steps,
                                  push_to_hub=False,
                                  log_level="error")

from transformers import Trainer
trainer = Trainer(model=model, args=training_args,
                  compute_metrics=compute_metrics,
                  train_dataset=emotions_encoded["train"],  # column_names显示为<class 'list'>: ['text', 'label', 'input_ids', 'attention_mask'] 
                  eval_dataset=emotions_encoded["validation"],
                  tokenizer=tokenizer)

trainer.train()

排查过程

  1. 检查trainer.py的get_train_dataloader函数,发现其中的train_dataset的column_names为['label', 'input_ids', 'attention_mask'],标签字段存在。
def get_train_dataloader(self) -> DataLoader:
        return DataLoader(
            train_dataset,
            batch_size=self._train_batch_size,
            sampler=train_sampler,
            collate_fn=data_collator,
            drop_last=self.args.dataloader_drop_last,
            num_workers=self.args.dataloader_num_workers,
            pin_memory=self.args.dataloader_pin_memory,
            worker_init_fn=seed_worker,
        )
  1. 追踪到datasets/arrow_dataset.py的_getitem函数,由于kwargs为None,代码使用了self._format_columns,而该值仅为['input_ids', 'attention_mask'],导致label字段被过滤。
def _getitem(self, key: Union[int, slice, str, ListLike[int]], **kwargs) -> Union[Dict, List]:
    format_columns = kwargs["format_columns"] if "format_columns" in kwargs else self._format_columns
  1. 最终在dataloder.py的_BaseDataLoaderIter.__next__函数中,返回的data仅包含input_ids和attention_mask,无label字段,导致模型无法计算loss。
def __next__(self) -> Any:
        with torch.autograd.profiler.record_function(self._profile_name):
            if self._sampler_iter is None:
                # TODO(https://github.com/pytorch/pytorch/issues/76750)
                self._reset()  # type: ignore[call-arg]
            data = self._next_data()

补充信息

模型定义

from transformers import AutoModelForSequenceClassification
model = (AutoModelForSequenceClassification
         .from_pretrained("distilbert-base-uncased", num_labels = 6)
         .to("cpu"))

版本信息

transformers                  4.30.2
datasets                      3.0.0
torch                         2.3.1+cpu
Python 3.10.2

解决方案

问题根源是数据集的_format_columns未包含label,导致加载数据时被过滤。在初始化Trainer之前,对训练和验证数据集执行set_format操作,明确指定需要保留的字段:

emotions_encoded["train"].set_format(type='torch', columns=['input_ids', 'attention_mask', 'label'])
emotions_encoded["validation"].set_format(type='torch', columns=['input_ids', 'attention_mask', 'label'])

这样数据集的_format_columns就会包含label字段,Dataloader加载时就能把标签传递给模型,模型即可正常计算loss。

内容的提问来源于stack exchange,提问作者Tiina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 23:10:09