You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何在TensorFlow中复现多项式回归无法实现过拟合?

问题分析与解决方案

你没法复现过拟合的问题,主要是几个细节没处理好,我帮你拆解一下:

1. 多项式阶数不对

你设置的order=15,但模型里的w是15个参数,对应从x^0到x^14,实际是14阶多项式,而Sklearn例子里是15阶(需要16个参数,覆盖x^0到x^15)。阶数不够的话,模型的复杂度不足以过拟合30个样本。

2. 特征尺度差异导致优化困难

x在[0,1]区间,x^15的数值会非常小(接近0),而x^0是1,不同幂次的特征尺度差异极大。这会让梯度下降优化器很难调整高次项的参数——低次项的梯度量级远高于高次项,导致高次项参数更新缓慢,没法学到足以过拟合的波动。

3. 优化器选择与训练策略问题

你用的是梯度下降(GradientDescentOptimizer),这种优化器对学习率非常敏感,在特征尺度不一致的场景下很容易陷入局部最优或者收敛缓慢。换成自适应学习率的优化器(比如Adam)会更稳定,能更快让模型拟合到噪声。


修正后的代码

我针对上面的问题调整了代码,你可以直接运行看看过拟合效果:

import tensorflow as tf
import numpy as np
import matplotlib.pyplot as plt

def true_fun(X):
    return np.cos(1.5 * np.pi * X)

# Generate dataset
n_samples = 30
np.random.seed(0)
x_train = np.sort(np.random.rand(n_samples)) # Draw from uniform distribution
y_train = true_fun(x_train) + np.random.randn(n_samples) * 0.1
x_test = np.linspace(0, 1, 100)
y_true = true_fun(x_test)

# Helper function
def run_dir(base_dir, dirname='run'):
    """Number log directories incrementally"""
    import os
    import re
    pattern = re.compile(dirname+'_(\d+)')
    try:
        previous_runs = os.listdir(base_dir)
    except FileNotFoundError:
        previous_runs = []
    run_number = 0
    for name in previous_runs:
        match = pattern.search(name)
        if match:
            number = int(match.group(1))
            if number > run_number:
                run_number = number
    run_number += 1
    logdir = os.path.join(base_dir, dirname + '_%02d' % run_number)
    return(logdir)

# Define the polynomial model - 修正:order是多项式的阶数,参数数量是order+1
def model(X, w):
    """Polynomial model
    param X: data
    param w: coefficients in the polynomial regression
    returns: Polynomial function Y(X, w)
    """
    terms = []
    # w的长度是order+1,对应x^0到x^order
    for i in range(len(w)):
        term = tf.multiply(w[i], tf.pow(X, i))
        terms.append(term)
    return(tf.add_n(terms))

# Create the computation graph
# 修正:15阶多项式需要16个参数
order = 15
tf.reset_default_graph()
X = tf.placeholder("float")
Y = tf.placeholder("float")
# 修正:参数数量是order+1
w = tf.Variable([0.0]*(order+1), name="parameters")
lambda_reg = tf.placeholder('float', shape=[])
learning_rate_ph = tf.placeholder('float', shape=[])

y_model = model(X, w)
loss = tf.div(tf.reduce_mean(tf.square(Y-y_model)), 2) # Square error
loss_rg = tf.multiply(lambda_reg, tf.reduce_sum(tf.square(w))) # L2 penalty
loss_total = tf.add(loss, loss_rg)

loss_hist1 = tf.summary.scalar('loss', loss)
loss_hist2 = tf.summary.scalar('loss_rg', loss_rg)
loss_hist3 = tf.summary.scalar('loss_total', loss_total)
summary = tf.summary.merge([loss_hist1, loss_hist2, loss_hist3])

# 修正:换成Adam优化器,自适应学习率更稳定
train_op = tf.train.AdamOptimizer(learning_rate_ph).minimize(loss_total)
init = tf.global_variables_initializer()

def train(sess, x_train, y_train, lambda_val=0, epochs=5000, learning_rate=0.01):
    feed_dict={X: x_train, Y: y_train, lambda_reg: lambda_val, learning_rate_ph: learning_rate}
    logdir = run_dir("logs/polynomial_regression2/")
    writer = tf.summary.FileWriter(logdir)
    sess.run(init)
    for epoch in range(epochs):
        _, summary_str = sess.run([train_op, summary], feed_dict=feed_dict)
        writer.add_summary(summary_str, global_step=epoch)
    final_cost, final_cost_rg, w_learned = sess.run([loss, loss_rg, w], feed_dict=feed_dict)
    return final_cost, final_cost_rg, w_learned

def plot_test(w_learned, x_test, x_train, y_train):
    y_learned = calculate_y(x_test, w_learned)
    plt.scatter(x_train, y_train, label="Training samples")
    plt.plot(x_test, y_true, label="True function", color="green")
    plt.plot(x_test, y_learned,'r', label="Learned function (15th order)")
    plt.ylabel('y')
    plt.xlabel('x')
    plt.legend()
    plt.ylim(-1.5, 1.5) # 限制y轴范围,更清晰看到过拟合的波动
    plt.show()

def calculate_y(x, w):
    y = 0
    for i in range(len(w)):
        y += w[i] * np.power(x, i)
    return y

sess = tf.Session()
# 修正:增加训练轮数到5000,用Adam优化器的话学习率0.01就够
final_cost, final_cost_rg, w_learned = train(sess, x_train, y_train, lambda_val=0, learning_rate=0.01, epochs=5000)
sess.close()
plot_test(w_learned, x_test, x_train, y_train)

额外说明

  • 如果你坚持要用梯度下降,需要先对特征进行标准化(比如把每个x^i缩放到均值0方差1),否则高次项的参数很难更新。
  • 训练轮数也要足够多,15阶模型需要更多迭代才能收敛到过拟合状态,Adam优化器比SGD快很多,所以我换成了它。

内容的提问来源于stack exchange,提问作者sjdh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:58:09