基于Keras绘制梯度下降曲线时theta值出现NaN的原因
问题:SGD训练中theta值出现NaN的原因及解决办法
我用Keras实现了基于加州住房数据集的代码,对比随机梯度下降、批量梯度下降和小批量梯度下降的效果,绘制theta1/theta2的变化轨迹。但当训练轮数(epochs)增加到10以上时,存储theta值的列表偶尔会出现NaN值,请问这是什么原因?
我的实现代码如下:
import numpy as np import matplotlib.pyplot as plt from keras.models import Sequential from keras.layers import Dense from keras.optimizers import SGD from sklearn.datasets import fetch_california_housing from sklearn.preprocessing import StandardScaler # Load California housing dataset california_housing = fetch_california_housing() X, y = california_housing.data, california_housing.target # Normalize features scaler = StandardScaler() X_normalized = scaler.fit_transform(X) def build_model(): model = Sequential([ Dense(64, activation="relu",input_shape=(8,)), Dense(64, activation="relu"), Dense(1) ]) #model.compile(optimizer="rmsprop", loss="mse", metrics=["mae"]) return model sgd_optimizer = SGD(lr=0.001) # Stochastic Gradient Descent minibatch_sgd_optimizer = SGD(lr=0.001) # Mini-batch Gradient Descent batch_sgd_optimizer = SGD(lr=0.001) # Batch Gradient Descent # Compile the model model=build_model() model.compile(loss='mse', optimizer=sgd_optimizer) theta1_sgd, theta2_sgd = [], [] theta1_minibatch_sgd, theta2_minibatch_sgd = [], [] theta1_batch_sgd, theta2_batch_sgd = [], [] # Function to perform gradient descent and store theta values def perform_gradient_descent(optimizer, batch_size=None): theta1_list, theta2_list = [], [] loss_history = [] for _ in range(5): # Number of epochs history=model.fit(X_normalized, y, epochs=1, batch_size=batch_size, verbose=0) loss_history.append(history.history['loss'][0]) weights = model.layers[0].get_weights()[0].flatten() # Get current theta values theta1_list.append(weights[0]) theta2_list.append(weights[1]) print(theta1_list," ",theta2_list) return loss_history,theta1_list, theta2_list # Perform gradient descent with different optimizers loss_sgd, theta1_sgd, theta2_sgd = perform_gradient_descent(sgd_optimizer, batch_size=1) # Stochastic Gradient Descent loss_minibatch_sgd, theta1_minibatch_sgd, theta2_minibatch_sgd = perform_gradient_descent(minibatch_sgd_optimizer, batch_size=32) # Mini-batch Gradient Descent loss_batch_sgd, theta1_batch_sgd, theta2_batch_sgd = perform_gradient_descent(batch_sgd_optimizer, batch_size=len(X_normalized)) # Batch Gradient Descent # Plotting the loss versus number of epochs plt.figure(figsize=(10, 6)) plt.plot(range(1, len(loss_sgd) + 1), loss_sgd, label='Stochastic Gradient Descent') plt.plot(range(1, len(loss_minibatch_sgd) + 1), loss_minibatch_sgd, label='Mini-batch Gradient Descent') plt.plot(range(1, len(loss_batch_sgd) + 1), loss_batch_sgd, label='Batch Gradient Descent') plt.xlabel('Number of Epochs') plt.ylabel('Loss') plt.title('Loss vs. Number of Epochs') plt.legend() plt.grid(True) plt.show() # Plotting the gradient descent trajectories plt.figure(figsize=(10, 6)) plt.plot(theta1_sgd, theta2_sgd, label='Stochastic Gradient Descent', marker='o') plt.plot(theta1_minibatch_sgd, theta2_minibatch_sgd, label='Mini-batch Gradient Descent', marker='s') plt.plot(theta1_batch_sgd, theta2_batch_sgd, label='Batch Gradient Descent', marker='x') plt.xlabel('Theta 1') plt.ylabel('Theta 2') plt.title('Gradient Descent Trajectories') #plt.xlim(-0.08, -0.05) # Set limit for Theta 1 #plt.ylim(0.02, 0.03) # Set limit for Theta 2 plt.legend() plt.grid(True) plt.show()
原因分析
- 梯度爆炸/数值溢出:固定学习率0.001在训练轮数增加后,可能遇到极端大的梯度(尤其是随机梯度下降用单样本,噪声极大),导致权重更新后数值急剧膨胀,超出浮点数范围变成NaN。
- 模型状态污染:三种梯度下降方法共用同一个
model实例,前一种训练的权重状态会被后一种继承。如果前序训练已经让权重出现异常,后续训练会放大问题,最终产生NaN。 - 损失计算溢出:当权重异常大时,模型预测值会变得极大,计算MSE损失时超出浮点数范围变成无穷大,反向传播的梯度也会变成NaN,进而导致权重更新为NaN。
- ReLU激活的潜在问题:ReLU神经元如果持续输出0,权重更新会停滞,但更可能是梯度爆炸引发的连锁反应导致NaN。
解决办法
- 独立训练模型:在
perform_gradient_descent函数内重新创建并编译模型,确保每种梯度下降从初始状态开始,避免状态污染。修改后的函数示例:def perform_gradient_descent(optimizer, batch_size=None): theta1_list, theta2_list = [], [] loss_history = [] # 每次训练都新建模型 model = build_model() model.compile(loss='mse', optimizer=optimizer) for _ in range(10): # 调整为10轮 history=model.fit(X_normalized, y, epochs=1, batch_size=batch_size, verbose=0) loss_history.append(history.history['loss'][0]) weights = model.layers[0].get_weights()[0].flatten() theta1_list.append(weights[0]) theta2_list.append(weights[1]) return loss_history,theta1_list, theta2_list - 调整学习率:降低学习率至
0.0001,或者使用带衰减的SGD:SGD(lr=0.001, decay=1e-6),减缓权重更新幅度。 - 添加梯度裁剪:编译时限制梯度范数,比如
optimizer=SGD(lr=0.001, clipnorm=1.0),防止梯度过大导致溢出。 - 归一化目标变量:对y值也做标准化处理,减少损失计算的数值范围:
训练时用from sklearn.preprocessing import StandardScaler y_scaler = StandardScaler() y_normalized = y_scaler.fit_transform(y.reshape(-1, 1))y_normalized代替原y值,绘制结果时再反归一化即可。
内容的提问来源于stack exchange,提问作者Little
相关产品推荐
相关产品推荐

