Keras自定义欧氏距离损失函数训练时出现NaN问题求助
Hey there, let's break down exactly why your Keras model is spitting out NaN losses and how to get it back on track.
Root Causes of the NaN Loss
Your setup has a few key mismatches and numerical stability issues that lead to this problem:
- Mismatched loss & output activation: You're using a
softmaxoutput layer (which produces probability distributions for classification) paired with an Euclidean distance loss. Euclidean loss is designed for regression tasks (predicting continuous values), not classification. This misalignment confuses the optimizer, leading to erratic gradient updates that can blow up into NaNs. - Numerical instability from
sqrt: WhileK.sqrt()seems harmless, when model parameters drift to extreme values (especially with a tiny batch size of 1), you can hit floating-point precision issues. Even worse, the gradient ofsqrt(x)when x approaches 0 is infinite—this can trigger massive parameter updates that break your model. - Extreme gradient noise with batch_size=1: Training on single samples creates huge gradient swings. Combine that with a loss function that can produce large gradients, and you're almost guaranteed to get parameter updates that push your model into invalid numerical territory.
Fixes to Try
Option 1: Use the Standard Loss for Classification (Most Recommended)
Since you're working on MNIST digit classification, categorical crossentropy is the gold-standard loss to pair with softmax. This is optimized for classification tasks, stable, and will eliminate the NaN issue immediately:
model.compile(loss='categorical_crossentropy', optimizer='adam', metrics=['accuracy'])
This change will also make your model train faster and perform better than trying to force Euclidean loss onto a classification task.
Option 2: Adapt Euclidean Loss for Your Use Case (If You Must)
If you have a specific reason to use Euclidean distance, here's how to make it work:
- Swap out the output activation: Replace
softmaxwith either no activation (for raw continuous outputs) orsigmoid—Euclidean loss works best with continuous values, not normalized probabilities:# Replace your output layer with this model.add(Dense(10, activation=None)) # Or activation='sigmoid' if you want 0-1 range - Remove the
sqrtfor stability: Squared Euclidean distance (L2 loss) has the same optimization direction as Euclidean distance but avoids the numerical instability of square roots:def euc_dist_keras(y_true, y_pred): # Drop the K.sqrt() call return K.mean(K.sum(K.square(y_true - y_pred), axis=-1, keepdims=True)) - Tame gradient noise: Increase your batch size (32 or 64 is standard) and lower the learning rate to prevent extreme parameter updates:
from keras.optimizers import Adam # Use a smaller learning rate model.compile(loss=euc_dist_keras, optimizer=Adam(learning_rate=0.0001), metrics=['accuracy']) # Train with a larger batch size train_history = model.fit(x=X_train4D_normalize, y=Y_trainOneHot, validation_split=0.2, epochs=10, batch_size=32, verbose=1)
Verify the Fix
After making these changes, re-run your training—you'll see the loss stays consistent (no more NaNs) and the model trains smoothly.
内容的提问来源于stack exchange,提问作者Sherlock Lo

