You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Huggingface BERT训练无损失计算,触发ValueError问题求助

解决Huggingface Trainer微调DistilBERT时的Loss返回错误

问题核心

使用Huggingface Trainer工具微调distilbert-base-uncased进行10分类任务时,触发ValueError,提示模型仅返回logits未返回损失值(loss),导致无法获取训练损失、F1、准确率等核心指标。尝试传入带隐藏状态和不带隐藏状态的编码数据集,问题均未解决。

完整报错栈

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-124-76d295da3120> in <module>
     24                  tokenizer=tokenizer)
     25 
---> 26 trainer.train();

/opt/conda/lib/python3.7/site-packages/transformers/trainer.py in train(self, resume_from_checkpoint, trial, ignore_keys_for_eval, **kwargs)
   1503             resume_from_checkpoint=resume_from_checkpoint,
   1504             trial=trial,
-> 1505             ignore_keys_for_eval=ignore_keys_for_eval,
   1506         )
   1507 

/opt/conda/lib/python3.7/site-packages/transformers/trainer.py in _inner_training_loop(self, batch_size, args, resume_from_checkpoint, trial, ignore_keys_for_eval)
   1747                         tr_loss_step = self.training_step(model, inputs)
   1748                 else:
-> 1749                     tr_loss_step = self.training_step(model, inputs)
   1750 
   1751                 if (

/opt/conda/lib/python3.7/site-packages/transformers/trainer.py in training_step(self, model, inputs)
   2506 
   2507         with self.compute_loss_context_manager():
-> 2508             loss = self.compute_loss(model, inputs)
   2509 
   2510         if self.args.n_gpu > 1:

/opt/conda/lib/python3.7/site-packages/transformers/trainer.py in compute_loss(self, model, inputs, return_outputs)
   2552             if isinstance(outputs, dict) and "loss" not in outputs:
   2553                 raise ValueError(
-> 2554                     "The model did not return a loss from the inputs, only the following keys: "
   2555                     f"{','.join(outputs.keys())}. For reference, the inputs it received are {','.join(inputs.keys())}."
   2556                 )

ValueError: The model did not return a loss from the inputs, only the following keys: logits. For reference, the inputs it received are input_ids,attention_mask.

问题排查

从代码和样本数据中发现3个关键问题:

  1. 标签列名不匹配:数据集中标签列名为Primary Label(带空格),但Huggingface Trainer默认期望标签列名为labels。模型计算损失时找不到标签输入,因此仅返回logits,无法生成loss。
  2. 指标计算函数笔误:compute_metrics函数中,acc = accuracy_score(labels, preds)里的preds未定义,正确变量名应为pred(前面已定义pred = pred.predictions.argmax(-1))。
  3. 冗余的隐藏状态提取:中间执行的extract_hidden_states步骤对分类微调无意义,反而可能干扰数据集结构,可直接移除。

修复方案

1. 重命名标签列

将数据集的Primary Label列重命名为labels,确保模型能识别标签输入:

# 在数据集加载后添加重命名操作
cd = cd.rename_column("Primary Label", "labels")

2. 修正指标计算函数

修复变量名错误:

from sklearn.metrics import accuracy_score, f1_score
def compute_metrics(pred):
    labels = pred.label_ids
    preds = pred.predictions.argmax(-1)  # 统一变量名为preds
    f1 = f1_score(labels, preds, average="weighted")
    acc = accuracy_score(labels, preds)
    return {"accuracy": acc, "f1": f1}

3. 移除冗余步骤

删除extract_hidden_states相关代码,分类微调无需提前提取隐藏状态,模型会自动处理。

修正后的完整代码

# 1. 加载并处理数据集
category_data = load_dataset("csv", data_files="testdatafinal.csv")
category_data = category_data.remove_columns(["someid", "someid", "somedimension"])
category_data = category_data['train']
train_testvalid = category_data.train_test_split(test_size=0.3)
test_valid = train_testvalid['test'].train_test_split(test_size=0.5)
from datasets.dataset_dict import DatasetDict
cd = DatasetDict({
    'train': train_testvalid['train'],
    'test': test_valid['test'],
    'valid': test_valid['train']})

# 重命名标签列
cd = cd.rename_column("Primary Label", "labels")

# 2. 文本编码(假设tokenizer已定义)
def preprocess_function(examples):
    return tokenizer(examples["Transcript"], truncation=True, padding=True)

transcripts_encoded = cd.map(preprocess_function, batched=True)
transcripts_encoded_one = transcripts_encoded.set_format("torch",
                              columns=["input_ids", "attention_mask", "labels"])

# 3. 定义分类模型
model_checkpoint = 'distilbert-base-uncased'
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
from transformers import AutoModelForSequenceClassification
num_labels = 10
model =(AutoModelForSequenceClassification
       .from_pretrained(model_checkpoint, num_labels=num_labels)
       .to(device))

# 4. 设置训练参数与Trainer
from transformers import Trainer, TrainingArguments
batch_size = 10
logging_steps = len(transcripts_encoded_one["train"]) // batch_size
model_name = f"{model_checkpoint}-finetuned-transcripts"
training_args = TrainingArguments(output_dir=model_name,
                                 num_train_epochs=2,
                                 learning_rate=2e-5,
                                 per_device_train_batch_size=batch_size,
                                 per_device_eval_batch_size=batch_size,
                                 weight_decay=0.01,
                                 evaluation_strategy="epoch",
                                 disable_tqdm=False,
                                 logging_steps=logging_steps,
                                 push_to_hub=False,
                                 log_level="error")

trainer = Trainer(model=model, args=training_args,
                 compute_metrics=compute_metrics,
                 train_dataset=transcripts_encoded_one["train"],
                 eval_dataset=transcripts_encoded_one["valid"],
                 tokenizer=tokenizer)

# 启动训练
trainer.train();

验证

运行修正后的代码后,检查训练样本结构:

print(trainer.train_dataset[0])
# 应输出类似:
# {'labels': 0,  # 已转换为数字ID(tokenizer编码时自动处理)
#  'input_ids': tensor([...]),
#  'attention_mask': tensor([...])}

此时训练会正常返回loss值,并在每个epoch结束后输出准确率和F1指标。

内容的提问来源于stack exchange,提问作者Wesson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 02:55:45