You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Keras绘制梯度下降曲线时theta值出现NaN的原因

问题:SGD训练中theta值出现NaN的原因及解决办法

我用Keras实现了基于加州住房数据集的代码,对比随机梯度下降、批量梯度下降和小批量梯度下降的效果,绘制theta1/theta2的变化轨迹。但当训练轮数(epochs)增加到10以上时,存储theta值的列表偶尔会出现NaN值,请问这是什么原因?

我的实现代码如下:

import numpy as np
import matplotlib.pyplot as plt
from keras.models import Sequential
from keras.layers import Dense
from keras.optimizers import SGD
from sklearn.datasets import fetch_california_housing
from sklearn.preprocessing import StandardScaler

# Load California housing dataset
california_housing = fetch_california_housing()
X, y = california_housing.data, california_housing.target

# Normalize features
scaler = StandardScaler()
X_normalized = scaler.fit_transform(X)

def build_model():
    model = Sequential([
    Dense(64, activation="relu",input_shape=(8,)),
    Dense(64, activation="relu"),
    Dense(1)
    ])
    #model.compile(optimizer="rmsprop", loss="mse", metrics=["mae"])
    return model


sgd_optimizer = SGD(lr=0.001)  # Stochastic Gradient Descent
minibatch_sgd_optimizer = SGD(lr=0.001)  # Mini-batch Gradient Descent
batch_sgd_optimizer = SGD(lr=0.001)  # Batch Gradient Descent


# Compile the model
model=build_model()
model.compile(loss='mse', optimizer=sgd_optimizer)

theta1_sgd, theta2_sgd = [], []
theta1_minibatch_sgd, theta2_minibatch_sgd = [], []
theta1_batch_sgd, theta2_batch_sgd = [], []


# Function to perform gradient descent and store theta values
def perform_gradient_descent(optimizer, batch_size=None):
    theta1_list, theta2_list = [], []
    loss_history = []
    for _ in range(5):  # Number of epochs
        history=model.fit(X_normalized, y, epochs=1, batch_size=batch_size, verbose=0)
        loss_history.append(history.history['loss'][0])
        weights = model.layers[0].get_weights()[0].flatten()  # Get current theta values
        theta1_list.append(weights[0])
        theta2_list.append(weights[1])
    print(theta1_list," ",theta2_list)
    return loss_history,theta1_list, theta2_list

# Perform gradient descent with different optimizers
loss_sgd, theta1_sgd, theta2_sgd = perform_gradient_descent(sgd_optimizer, batch_size=1)  # Stochastic Gradient Descent
loss_minibatch_sgd, theta1_minibatch_sgd, theta2_minibatch_sgd = perform_gradient_descent(minibatch_sgd_optimizer, batch_size=32)  # Mini-batch Gradient Descent
loss_batch_sgd, theta1_batch_sgd, theta2_batch_sgd = perform_gradient_descent(batch_sgd_optimizer, batch_size=len(X_normalized))  # Batch Gradient Descent

# Plotting the loss versus number of epochs
plt.figure(figsize=(10, 6))
plt.plot(range(1, len(loss_sgd) + 1), loss_sgd, label='Stochastic Gradient Descent')
plt.plot(range(1, len(loss_minibatch_sgd) + 1), loss_minibatch_sgd, label='Mini-batch Gradient Descent')
plt.plot(range(1, len(loss_batch_sgd) + 1), loss_batch_sgd, label='Batch Gradient Descent')
plt.xlabel('Number of Epochs')
plt.ylabel('Loss')
plt.title('Loss vs. Number of Epochs')
plt.legend()
plt.grid(True)
plt.show()

# Plotting the gradient descent trajectories
plt.figure(figsize=(10, 6))
plt.plot(theta1_sgd, theta2_sgd, label='Stochastic Gradient Descent', marker='o')
plt.plot(theta1_minibatch_sgd, theta2_minibatch_sgd, label='Mini-batch Gradient Descent', marker='s')
plt.plot(theta1_batch_sgd, theta2_batch_sgd, label='Batch Gradient Descent', marker='x')
plt.xlabel('Theta 1')
plt.ylabel('Theta 2')
plt.title('Gradient Descent Trajectories')
#plt.xlim(-0.08, -0.05)  # Set limit for Theta 1
#plt.ylim(0.02, 0.03)  # Set limit for Theta 2
plt.legend()
plt.grid(True)
plt.show()

原因分析

  • 梯度爆炸/数值溢出:固定学习率0.001在训练轮数增加后,可能遇到极端大的梯度(尤其是随机梯度下降用单样本,噪声极大),导致权重更新后数值急剧膨胀,超出浮点数范围变成NaN。
  • 模型状态污染:三种梯度下降方法共用同一个model实例,前一种训练的权重状态会被后一种继承。如果前序训练已经让权重出现异常,后续训练会放大问题,最终产生NaN。
  • 损失计算溢出:当权重异常大时,模型预测值会变得极大,计算MSE损失时超出浮点数范围变成无穷大,反向传播的梯度也会变成NaN,进而导致权重更新为NaN。
  • ReLU激活的潜在问题:ReLU神经元如果持续输出0,权重更新会停滞,但更可能是梯度爆炸引发的连锁反应导致NaN。

解决办法

  • 独立训练模型:在perform_gradient_descent函数内重新创建并编译模型,确保每种梯度下降从初始状态开始,避免状态污染。修改后的函数示例:
    def perform_gradient_descent(optimizer, batch_size=None):
        theta1_list, theta2_list = [], []
        loss_history = []
        # 每次训练都新建模型
        model = build_model()
        model.compile(loss='mse', optimizer=optimizer)
        for _ in range(10):  # 调整为10轮
            history=model.fit(X_normalized, y, epochs=1, batch_size=batch_size, verbose=0)
            loss_history.append(history.history['loss'][0])
            weights = model.layers[0].get_weights()[0].flatten()
            theta1_list.append(weights[0])
            theta2_list.append(weights[1])
        return loss_history,theta1_list, theta2_list
    
  • 调整学习率:降低学习率至0.0001,或者使用带衰减的SGD:SGD(lr=0.001, decay=1e-6),减缓权重更新幅度。
  • 添加梯度裁剪:编译时限制梯度范数,比如optimizer=SGD(lr=0.001, clipnorm=1.0),防止梯度过大导致溢出。
  • 归一化目标变量:对y值也做标准化处理,减少损失计算的数值范围:
    from sklearn.preprocessing import StandardScaler
    y_scaler = StandardScaler()
    y_normalized = y_scaler.fit_transform(y.reshape(-1, 1))
    
    训练时用y_normalized代替原y值,绘制结果时再反归一化即可。

内容的提问来源于stack exchange,提问作者Little

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 00:42:03