为何增大训练集时TensorFlow模型损失值趋于无穷大?
问题描述
构建了一个简单的TensorFlow线性模型,用于拟合y=2x-1。当使用19个训练样本时,模型训练正常,损失趋近于0,预测结果准确;但仅新增1个样本(共20个)后,训练过程中损失值趋于无穷大,模型完全失效。两个模型结构完全一致,仅训练样本数量相差1个。
复现代码如下:
import tensorflow as tf import numpy as np from tensorflow import keras print(tf.__version__) def hw_function(x): y = (2. * x) - 1. return y # Build a simple Sequential model model = tf.keras.Sequential([ tf.keras.Input(shape=(1,)), tf.keras.layers.Dense(units=1)]) # Compile the model model.compile(optimizer='sgd', loss='mean_squared_error') # Declare model inputs and outputs for training xs=[x for x in range(-1, 19, 1)] ys=[x for x in range(-3, 36, 2)] xs=np.array(xs, dtype=float) ys=np.array(ys, dtype=float) # Train the model model.fit(xs, ys, verbose=1, epochs=500) # Make a prediction p = np.array([100.0, 900.0], dtype=float) print(model.predict(p)) # Build exactly the same model but have one more training example model2 = tf.keras.Sequential([ tf.keras.Input(shape=(1,)), tf.keras.layers.Dense(units=1)]) model2.compile(optimizer='sgd', loss='mean_squared_error') xs2=[x for x in range(-1, 18, 1)] ys2=[x for x in range(-3, 34, 2)] xs2=np.array(xs2, dtype=float) ys2=np.array(ys2, dtype=float) # Train the model model2.fit(xs2, ys2, verbose=1, epochs=500) p = np.array([100.0, 900.0], dtype=float) print(model2.predict(p))
问题根源
核心原因是SGD优化器的默认学习率(0.01)相对于输入数据的尺度偏大,样本数量变化后梯度更新的稳定性被打破:
- 输入
xs的范围是[-1, 18],数值跨度较大,初始随机权重下第一次梯度计算的幅度可能很高; - 当样本数量增加到20时,全batch训练的梯度虽为均值,但某次更新步长过大导致权重偏离最优值,后续迭代中损失持续放大,最终趋于无穷大;
- 19个样本时恰好初始梯度和更新步长处于相对稳定范围,因此能收敛到正确结果。
解决办法
1. 标准化输入数据
将输入缩放到均值为0、方差为1的范围,消除数据尺度对梯度的影响:
# 对xs和xs2做标准化处理 xs = (xs - np.mean(xs)) / np.std(xs) xs2 = (xs2 - np.mean(xs2)) / np.std(xs2)
2. 降低SGD学习率
调小学习率,减小每次权重更新的步长,避免震荡发散:
model.compile(optimizer=tf.keras.optimizers.SGD(learning_rate=0.001), loss='mean_squared_error') model2.compile(optimizer=tf.keras.optimizers.SGD(learning_rate=0.001), loss='mean_squared_error')
3. 使用自适应优化器(推荐)
改用Adam等自适应学习率的优化器,它会根据梯度自动调整学习率,稳定性远高于SGD:
model.compile(optimizer='adam', loss='mean_squared_error') model2.compile(optimizer='adam', loss='mean_squared_error')
内容的提问来源于stack exchange,提问作者Fred Myers
相关产品推荐
相关产品推荐

