You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Seq2Seq Trainer训练时inf/nan异常问题求助

微调Seq2Seq模型时Embedding层出现inf/nan问题排查与解决

在Shell Code数据集上使用Hugging Face Trainer微调Seq2Seq模型时,训练启动即触发inf/nan错误,报错指向Embedding层,无法继续训练。

训练代码

from transformers import PreTrainedTokenizerFast

tokenizer = PreTrainedTokenizerFast(tokenizer_file="tkn1.json", padding_side="right") 
special_tokens={'pad_token': "[PAD]"}

tokenizer.add_special_tokens(special_tokens)

#  token_wrap = PreTrainedTokenizer()
data_collator = DataCollatorForSeq2Seq(tokenizer=tokenizer, model=model)

training_args = Seq2SeqTrainingArguments(
    output_dir="./results",
    evaluation_strategy="epoch",
    lr_scheduler_type = "cosine",
    weight_decay=0.01,
    save_total_limit=3,
    per_device_train_batch_size=128,
    num_train_epochs=5,
    warmup_ratio=0.06,
    learning_rate=1.0e-04,
    # fp16=True,
    debug=["underflow_overflow"]
)

trainer = Seq2SeqTrainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_datasets["test"],
    eval_dataset=tokenized_datasets["test"],
    tokenizer=tokenizer,
    data_collator=data_collator,
)
# trainer.train()
# print(tokenizer.)
trainer.train()
# eval_loss = trainer.evaluate()
# print(f">>> Perplexity: {math.exp(eval_loss['eval_loss']):.2f}")

运行报错输出

You're using a PreTrainedTokenizerFast tokenizer. Please note that with a fast tokenizer, using the `__call__` method is faster than using a method to encode the text followed by a call to the `pad` method to get a padded encoding.


Detected inf/nan during batch_number=0
Last 1 forward frames:
abs min  abs max  metadata
                  shared Embedding
5.42e-06 2.04e+04 weight
0.00e+00 1.46e+03 input[0]
1.56e-03 2.04e+04 output


---------------------------------------------------------------------------

ValueError                                Traceback (most recent call last)

<ipython-input-120-ff4a54906908> in <module>
     33 # trainer.train()
     34 # print(tokenizer.)
---&gt; 35 trainer.train()
     36 # eval_loss = trainer.evaluate()
     37 # print(f">>> Perplexity: {math.exp(eval_loss['eval_loss']):.2f}")

9 frames

/usr/local/lib/python3.8/dist-packages/transformers/debug_utils.py in forward_hook(self, module, input, output)
    278 
    279             # now we can abort, as it's pointless to continue running
--&gt; 280             raise ValueError(
    281                 "DebugUnderflowOverflow: inf/nan detected, aborting as there is no point running further. "
    282                 "Please scroll up above this traceback to see the activation values prior to this event."

ValueError: DebugUnderflowOverflow: inf/nan detected, aborting as there is no point running further. Please scroll up above this traceback to see the activation values prior to this event.

解决建议

  • 检查Embedding层权重:报错显示Embedding权重的绝对最大值达到2.04e+04,数值异常偏高。可能是模型加载时权重损坏,或自定义模型的Embedding初始化方式不当。可以尝试重新加载预训练模型,或手动用Xavier/He初始化重置Embedding层权重。
  • 缩小批次大小:当前per_device_train_batch_size=128过大,容易引发梯度爆炸。先降至16或32,再逐步调整到合适值。
  • 降低学习率:将learning_rate从1e-4降至5e-5或1e-5,减少权重更新的步长,避免数值溢出。
  • 添加梯度裁剪:在Seq2SeqTrainingArguments中加入gradient_clipping=1.0,限制梯度的最大范数,防止梯度爆炸导致权重异常。
  • 排查数据集合法性:抽样检查tokenized_datasets的input_ids和labels,确认是否存在超长序列、非法令牌或NaN值,确保数据预处理无问题。
  • 确认特殊令牌配置:检查pad_token是否被模型正确识别,确保模型在计算时能正确忽略padding部分,避免无效计算引发的数值异常。
  • 关闭混合精度:确保fp16=False(代码中已注释,但需确认实际运行状态),混合精度训练可能在部分场景下引发数值不稳定。

内容的提问来源于stack exchange,提问作者Gitanjali Mannepalli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 16:15:56