You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Huggingface Trainer API微调GPT2?解决无Loss返回错误

问题:微调GPT2时Trainer API返回无loss错误

我是机器学习新手,正在学习Huggingface Trainer API和Transformers库。最终目标是在自定义数据集上微调GODEL(或比DialoGPT更优的模型,目前已跑通DialoGPT的自定义训练),想用Trainer API实现(如果思路错了请指正)。先拿GPT2做基础测试,修改BERT文本分类的教程代码适配文本生成后,调用trainer.train()时出现如下错误:

ValueError: The model did not return a loss from the inputs, only the following keys: logits,past_key_values. For reference, the inputs it received are input_ids,attention_mask.

不确定问题出在metric设置、Trainer参数、compute_metrics还是其他环节,附上完整代码:

from datasets import load_dataset
from transformers import AutoTokenizer
from transformers import AutoModelForCausalLM
import numpy as np
import evaluate
from transformers import TrainingArguments, Trainer
import torch
from pynvml import *


training_args = TrainingArguments(output_dir='test_trainer', 
                                  evaluation_strategy='epoch',
                                  per_device_train_batch_size=1,
                                  per_device_eval_batch_size=1,
                                  gradient_accumulation_steps=20, # I'm paranoid about memory
                                  num_train_epochs = 2,
                                  fp16=False,)

# Load model and specify number of labels that text can be classified as
model = AutoModelForCausalLM.from_pretrained("gpt2").to("cuda")
tokenizer = AutoTokenizer.from_pretrained("gpt2")
metric = evaluate.load("accuracy")


tokenizer.pad_token = tokenizer.eos_token

# This might be where the problem is, but I'm not sure how to write it for straight-up simple text generation
def compute_metrics(eval_pred):
    logits, labels = eval_pred
    predictions = np.argmax(logits, axis=-1) # Get largest logit / the prediction
    return metric.compute(predictions=predictions
                          )


def tokenize_function(examples):
    return tokenizer(examples["text"],
                    #  padding="max_length",
                    padding=True,
                     truncation=True)


# Load a smaller, sampled dataset instead of the full one so that it doesn't take 10 years to run this code each time
dataset = load_dataset('codyburker/yelp_review_sampled')

dataset = dataset.map(tokenize_function,batched=True,
                        batch_size=1)

# Debug print statements to see what each object looks like
# print(dataset["train"][100])
# print(dataset["test"][100])

small_train_dataset = dataset["train"].shuffle(seed=42).select(range(1000))
small_eval_dataset = dataset["test"].shuffle(seed=42).select(range(1000))


trainer = Trainer(
    model=model,
    args = training_args,
    train_dataset=small_train_dataset,
    # eval_dataset=small_eval_dataset,
    # compute_metrics=compute_metrics, #commented out because the compute_metrics is unchanged from the original text classification code
)

trainer.train()

解决方案

核心原因

GPT2这类因果语言模型的训练需要标签输入来计算自回归损失,但你的数据集里没有labels字段,模型无法自动生成损失值。

1. 修改tokenize函数,添加labels字段

在tokenize时,把input_ids复制为labels,并将padding位置替换为-100(模型会自动忽略这些位置的损失计算):

def tokenize_function(examples):
    tokenized = tokenizer(examples["text"], padding=True, truncation=True)
    # 生成labels,padding位用-100标记
    tokenized["labels"] = [
        [-100 if token == tokenizer.pad_token_id else token for token in seq]
        for seq in tokenized["input_ids"]
    ]
    return tokenized

2. 关于评估指标的说明

文本生成任务不适合用分类任务的accuracy指标:

  • 如果只需要基础评估,可以直接依赖Trainer自动计算的困惑度(perplexity),不需要自定义compute_metrics
  • 如果需要更针对性的评估,可以换成生成任务专用指标(如BLEU、ROUGE),但需要额外实现对应的计算逻辑

3. 其他可选优化

  • 你的gradient_accumulation_steps=20搭配per_device_train_batch_size=1等效于总batch size 20,若显存允许,可适当调大per_device_train_batch_size以减少累加步数,提升训练效率
  • 若后续需要评估生成效果,可在TrainingArguments中设置predict_with_generate=True,并补充生成参数(如max_new_tokens)

内容的提问来源于stack exchange,提问作者Evan Armstrong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 03:37:03