You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GPT2损失函数理解疑问:如何构造标签使损失趋近于零?

GPT2损失计算的困惑解析

在理解GPT2的损失计算时遇到问题:希望通过设置包含预期生成目标的标签,让模型损失趋近于0。现有输入文本input_text = "Welcome to New York",模型当前预测下一个词是City,但无论把input_text设为标签,还是构造标签为"Welcome to New York City",损失都不为0。


情况1:输入带EOS标签,目标为完整句子

测试代码

from transformers import GPT2LMHeadModel, GPT2Tokenizer

model_name = 'gpt2'
tokenizer = GPT2Tokenizer.from_pretrained(model_name,model_max_length=1024,padding_side='left')
tokenizer.pad_token = tokenizer.eos_token # == <|endoftext|> = 50256
model = GPT2LMHeadModel.from_pretrained(model_name)

batch_size=5
input_text  = "<|endoftext|> Welcome to New York"
target_text = "Welcome to New York City"

# 编码输入
encoding = tokenizer(input_text,padding=True,max_length=batch_size,truncation=True,return_tensors="pt",)
input_ids, attention_mask = encoding.input_ids, encoding.attention_mask
# 编码目标文本
target_encoding = tokenizer(target_text,padding=True, max_length=batch_size, truncation=True,return_tensors="pt",)
labels = target_encoding.input_ids
# 将标签中的padding token替换为-100,避免被损失计算纳入
labels[labels == tokenizer.pad_token_id] = -100  # 本次无padding
print(f"input_ids={input_ids}")
print(f"attention_mask={attention_mask}") # 全为1
print(f"labels ={labels}")
# 前向传播
outputs = model(input_ids=input_ids,labels=labels) 
print(f"Model Loss {outputs.loss}")
# 测试模型预测下一个词
outputs = model.generate(input_ids=input_ids, attention_mask=attention_mask,max_new_tokens=1)
answer = tokenizer.decode(outputs[0], skip_special_tokens=False)
print(f"Result '{answer}'")

输出结果

input_ids=tensor([[50256, 19134,   284,   968,  1971]]) # 不清楚输入中的EOS token(50256)对模型的影响
attention_mask=tensor([[1, 1, 1, 1, 1]])
labels =tensor([[14618,   284,   968,  1971,  2254]]) # 2254对应City,是模型应该预测的词
Model Loss 8.248174667358398
Setting `pad_token_id` to `eos_token_id`:50256 for open-end generation.
Result '<|endoftext|> Welcome to New York City'

情况2:输入与目标文本相同

测试代码片段

input_text  = "Welcome to New York"
target_text = input_text

输出结果

input_ids=tensor([[14618,   284,   968,  1971]]) # 1971对应York
attention_mask=tensor([[1, 1, 1, 1]])
labels =tensor([[14618,   284,   968,  1971]])
Model Loss 3.2614505290985107
Setting `pad_token_id` to `eos_token_id`:50256 for open-end generation.
Result 'Welcome to New York City'

损失不为0的核心原因

  1. 预训练模型的固有特性:你使用的是预训练完成的GPT2,它的权重基于海量通用语料训练,不会仅通过一次前向传播就适配特定样本的标签。只有针对该样本进行微调训练,迭代更新权重后,损失才会逐步趋近于0。

  2. 标签左移的计算逻辑:GPT2传入labels时,会自动对标签左移——用输入的第i个token预测标签的第i+1个token,且不会计算输入最后一个token的损失(无对应下一个标签token)。

    • 第一种情况中,输入序列是[EOS, Welcome, to, New, York],标签序列是[Welcome, to, New, York, City]。实际计算损失的是:用EOS预测Welcome、用Welcome预测to、用to预测New、用New预测York这四个任务,你期望的“用York预测City”并没有被纳入损失计算(输入最后一个token无对应标签的下一位)。
    • 第二种情况中,输入和标签都是[Welcome, to, New, York],损失计算的是:用Welcome预测to、用to预测New、用New预测York这三个任务,预训练模型未针对该样本优化,因此损失不为0。

让损失趋近于0的方法

  • 针对样本微调:冻结模型部分层或全量微调,用该样本的输入和正确标签进行多轮迭代训练,模型权重会逐步调整,损失会逐渐降低。
  • 匹配正确的输入-标签对应关系:如果只想让模型学会用Welcome to New York预测City,可设置输入为"Welcome to New York",标签为"to New York City"(符合左移后的对应逻辑),或在计算损失时仅关注最后一个token的预测结果(即City)。

内容的提问来源于stack exchange,提问作者Alex Punnen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 23:17:14