You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何增大训练集时TensorFlow模型损失值趋于无穷大?

问题描述

构建了一个简单的TensorFlow线性模型,用于拟合y=2x-1。当使用19个训练样本时,模型训练正常,损失趋近于0,预测结果准确;但仅新增1个样本(共20个)后,训练过程中损失值趋于无穷大,模型完全失效。两个模型结构完全一致,仅训练样本数量相差1个。

复现代码如下:

import tensorflow as tf
import numpy as np
from tensorflow import keras

print(tf.__version__)

def hw_function(x):
    y = (2. * x) - 1.
    return y

# Build a simple Sequential model
model = tf.keras.Sequential([
    tf.keras.Input(shape=(1,)),
    tf.keras.layers.Dense(units=1)])

# Compile the model
model.compile(optimizer='sgd', loss='mean_squared_error')

# Declare model inputs and outputs for training
xs=[x for x in range(-1, 19, 1)]
ys=[x for x in range(-3, 36, 2)]

xs=np.array(xs, dtype=float)
ys=np.array(ys, dtype=float)

# Train the model
model.fit(xs, ys, verbose=1, epochs=500)

# Make a prediction
p = np.array([100.0, 900.0], dtype=float)
print(model.predict(p))


# Build exactly the same model but have one more training example
model2 = tf.keras.Sequential([
    tf.keras.Input(shape=(1,)),
    tf.keras.layers.Dense(units=1)])
model2.compile(optimizer='sgd', loss='mean_squared_error')
xs2=[x for x in range(-1, 18, 1)]
ys2=[x for x in range(-3, 34, 2)]

xs2=np.array(xs2, dtype=float)
ys2=np.array(ys2, dtype=float)

# Train the model
model2.fit(xs2, ys2, verbose=1, epochs=500)
p = np.array([100.0, 900.0], dtype=float)
print(model2.predict(p))

问题根源

核心原因是SGD优化器的默认学习率(0.01)相对于输入数据的尺度偏大,样本数量变化后梯度更新的稳定性被打破:

  • 输入xs的范围是[-1, 18],数值跨度较大,初始随机权重下第一次梯度计算的幅度可能很高;
  • 当样本数量增加到20时,全batch训练的梯度虽为均值,但某次更新步长过大导致权重偏离最优值,后续迭代中损失持续放大,最终趋于无穷大;
  • 19个样本时恰好初始梯度和更新步长处于相对稳定范围,因此能收敛到正确结果。

解决办法

1. 标准化输入数据

将输入缩放到均值为0、方差为1的范围,消除数据尺度对梯度的影响:

# 对xs和xs2做标准化处理
xs = (xs - np.mean(xs)) / np.std(xs)
xs2 = (xs2 - np.mean(xs2)) / np.std(xs2)

2. 降低SGD学习率

调小学习率,减小每次权重更新的步长,避免震荡发散:

model.compile(optimizer=tf.keras.optimizers.SGD(learning_rate=0.001), loss='mean_squared_error')
model2.compile(optimizer=tf.keras.optimizers.SGD(learning_rate=0.001), loss='mean_squared_error')

3. 使用自适应优化器(推荐)

改用Adam等自适应学习率的优化器,它会根据梯度自动调整学习率,稳定性远高于SGD:

model.compile(optimizer='adam', loss='mean_squared_error')
model2.compile(optimizer='adam', loss='mean_squared_error')

内容的提问来源于stack exchange,提问作者Fred Myers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 22:48:25