You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

VAE训练时总损失与KL损失变为NaN的问题求助

问题:VAE训练中总损失与KL损失出现NaN

我正在基于黑白猫图像数据集训练VAE,已完成以下排查:

  • 以255为最大像素值做归一化,所有像素值处于[0,1]区间,适配sigmoid激活函数;
  • 学习率设为0.0001,低于常用的0.0005;
  • 确认输入数据无NaN值;
  • 尝试用0.00001、0.0001、0.001、0.01等系数缩放KL损失以降低数值。

但训练时总损失与KL损失仍变为NaN,以下是损失函数代码:

def _calculate_reconstruction_loss(y_target, y_predicted):
    error = y_target - y_predicted
    reconstruction_loss = K.mean(K.square(error), axis=[1, 2, 3])
    return reconstruction_loss


def calculate_kl_loss(model):
    # wrap `_calculate_kl_loss` such that it takes the model as an argument,
    # returns a function which can take arbitrary number of arguments
    # (for compatibility with `metrics` and utility in the loss function)
    # and returns the kl loss
    def _calculate_kl_loss(*args):
        kl_loss = -0.5 * K.mean(1 + model.log_variance*0.01 - K.square(model.mu) - K.exp(model.log_variance*0.01), axis=1)
        return kl_loss
    return _calculate_kl_loss
解决建议
  • 限制log_variance的取值范围:KL损失中的K.exp(model.log_variance*0.01)如果log_variance过大,指数运算会导致数值爆炸产生NaN。可以在模型输出log_variance时添加约束,比如用K.clip(model.log_variance, min_value=-10, max_value=2),或用tanh激活后再缩放,避免指数后数值溢出。
  • 检查KL损失计算的数值稳定性:单独打印model.mu、model.log_variance的取值范围,确认是否存在mu平方过大、log_variance极端正负等异常情况,定位数值波动的源头。
  • 改用更稳定的KL损失计算方式:对exp的输入做截断,比如将K.exp(model.log_variance*0.01)改为K.exp(K.clip(model.log_variance*0.01, max_value=10)),防止指数运算产生无穷大值。
  • 添加梯度裁剪:在优化器中设置梯度裁剪,比如tf.keras.optimizers.Adam(learning_rate=1e-4, clipnorm=1.0),限制梯度的最大范数,避免参数更新时因梯度爆炸出现NaN。
  • 监控重构损失的数值范围:打印y_predicted的最大值和最小值,确认sigmoid输出是否稳定在[0,1]区间,极端值可能导致重构损失异常,间接引发总损失NaN。
  • 逐批次监控损失:训练时逐批次打印KL损失、重构损失的具体数值,定位首次出现NaN的批次,排查该批次数据或参数更新是否存在异常。

内容的提问来源于stack exchange,提问作者Aditya Shah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 19:44:55