You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras自定义欧氏距离损失函数训练时出现NaN问题求助

Why Your Custom Euclidean Loss is Causing NaN Values (and Fixes)

Hey there, let's break down exactly why your Keras model is spitting out NaN losses and how to get it back on track.

Root Causes of the NaN Loss

Your setup has a few key mismatches and numerical stability issues that lead to this problem:

  • Mismatched loss & output activation: You're using a softmax output layer (which produces probability distributions for classification) paired with an Euclidean distance loss. Euclidean loss is designed for regression tasks (predicting continuous values), not classification. This misalignment confuses the optimizer, leading to erratic gradient updates that can blow up into NaNs.
  • Numerical instability from sqrt: While K.sqrt() seems harmless, when model parameters drift to extreme values (especially with a tiny batch size of 1), you can hit floating-point precision issues. Even worse, the gradient of sqrt(x) when x approaches 0 is infinite—this can trigger massive parameter updates that break your model.
  • Extreme gradient noise with batch_size=1: Training on single samples creates huge gradient swings. Combine that with a loss function that can produce large gradients, and you're almost guaranteed to get parameter updates that push your model into invalid numerical territory.

Fixes to Try

Since you're working on MNIST digit classification, categorical crossentropy is the gold-standard loss to pair with softmax. This is optimized for classification tasks, stable, and will eliminate the NaN issue immediately:

model.compile(loss='categorical_crossentropy', optimizer='adam', metrics=['accuracy'])

This change will also make your model train faster and perform better than trying to force Euclidean loss onto a classification task.

Option 2: Adapt Euclidean Loss for Your Use Case (If You Must)

If you have a specific reason to use Euclidean distance, here's how to make it work:

  1. Swap out the output activation: Replace softmax with either no activation (for raw continuous outputs) or sigmoid—Euclidean loss works best with continuous values, not normalized probabilities:
    # Replace your output layer with this
    model.add(Dense(10, activation=None))  # Or activation='sigmoid' if you want 0-1 range
    
  2. Remove the sqrt for stability: Squared Euclidean distance (L2 loss) has the same optimization direction as Euclidean distance but avoids the numerical instability of square roots:
    def euc_dist_keras(y_true, y_pred):
        # Drop the K.sqrt() call
        return K.mean(K.sum(K.square(y_true - y_pred), axis=-1, keepdims=True))
    
  3. Tame gradient noise: Increase your batch size (32 or 64 is standard) and lower the learning rate to prevent extreme parameter updates:
    from keras.optimizers import Adam
    # Use a smaller learning rate
    model.compile(loss=euc_dist_keras, optimizer=Adam(learning_rate=0.0001), metrics=['accuracy'])
    # Train with a larger batch size
    train_history = model.fit(x=X_train4D_normalize, y=Y_trainOneHot, validation_split=0.2,
                              epochs=10, batch_size=32, verbose=1)
    

Verify the Fix

After making these changes, re-run your training—you'll see the loss stays consistent (no more NaNs) and the model trains smoothly.

内容的提问来源于stack exchange,提问作者Sherlock Lo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:28:10