手动实现多项式回归梯度下降与Sklearn模型系数不符问题求助
多项式回归手动梯度下降与sklearn结果不一致的问题
问题描述
手动实现多项式回归的梯度下降算法后,发现最终得到的参数始终和sklearn的LinearRegression输出不一致,即使调整学习率和迭代次数也无法解决。代码如下:
import numpy as np import pandas as pd from sklearn.linear_model import LinearRegression from matplotlib import pyplot as plt from sklearn.model_selection import train_test_split from sklearn.preprocessing import PolynomialFeatures # Creating Random X and Y Value np.random.seed(405) X = 6*np.random.rand(100,1)-3 Y = 0.8*(X**2) + 0.9*X + 2 + np.random.randn(100,1) Y = Y.reshape(-1,1) X_train,X_test,y_train,y_test = train_test_split(X,Y,test_size=0.2,random_state=2) poly = PolynomialFeatures(degree=2,include_bias=False) X_train_trans = poly.fit_transform(X_train) X_test_trans = poly.transform(X_test) model = LinearRegression() model.fit(X_train_trans,y_train) y_pred = model.predict(X_test_trans) print("sklearn LinearRegression参数:") print(f"截距: {model.intercept_}") print(f"系数: {model.coef_}") n = len(X_train_trans) β0 = 0 β1 = 1 β2 = 2 learning_rate = 0.0001 # Adjusted learning rate num_iterations = 50000 # Increased iterations n = len(X_train_trans) for iteration in range(num_iterations): # Compute predictions y_prediction = β0 + β1 * X_train_trans[:, 0] + β2 * X_train_trans[:, 1] # Compute the loss (Mean Squared Error) loss = (1/(2*n)) * np.sum((y_train - y_prediction)**2) # Compute gradients d_β0 = -(1 / n) * np.sum(y_train - y_prediction) d_β1 = -(1 / n) * np.sum((y_train - y_prediction) * X_train_trans[:, 0]) d_β2 = -(1 / n) * np.sum((y_train - y_prediction) * X_train_trans[:, 1]) # Update parameters β0 -= learning_rate * d_β0 β1 -= learning_rate * d_β1 β2 -= learning_rate * d_β2 # Print loss and parameters periodically if iteration % 1000 == 0: print(f"Iteration {iteration}: Loss = {loss}") print(f"\nFinal parameters (Gradient Descent): β0 = {β0}, β1 = {β1}, β2 = {β2}")
核心原因
- 特征未做标准化:
PolynomialFeatures生成的二次项特征(如X²)数值范围远大于一次项(X),梯度下降对不同尺度的特征收敛效率极低——大尺度特征的梯度更新幅度极小,导致参数无法逼近最优解。而sklearn的LinearRegression使用闭式解(最小二乘法),不受特征尺度影响,直接计算全局最优值。 - 学习率与迭代次数匹配问题:未做特征缩放时,即使调大迭代次数、调小学习率,也无法解决不同特征梯度更新幅度不均衡的问题,参数始终无法收敛到最优值。
解决方案
1. 添加特征标准化
使用StandardScaler对多项式变换后的特征进行缩放,让所有特征处于同一数值尺度。
2. 调整学习率和迭代次数
特征标准化后,可以使用更大的学习率(比如0.01),同时减少迭代次数(比如10000次即可收敛)。
3. 修正维度匹配问题
确保y_prediction的维度和y_train一致,避免广播错误。
修改后的代码
import numpy as np from sklearn.linear_model import LinearRegression from sklearn.model_selection import train_test_split from sklearn.preprocessing import PolynomialFeatures, StandardScaler # 创建数据集 np.random.seed(405) X = 6*np.random.rand(100,1)-3 Y = 0.8*(X**2) + 0.9*X + 2 + np.random.randn(100,1) Y = Y.reshape(-1,1) # 划分数据集 X_train,X_test,y_train,y_test = train_test_split(X,Y,test_size=0.2,random_state=2) # 多项式变换+特征标准化 poly = PolynomialFeatures(degree=2,include_bias=False) X_train_trans = poly.fit_transform(X_train) X_test_trans = poly.transform(X_test) scaler = StandardScaler() X_train_scaled = scaler.fit_transform(X_train_trans) X_test_scaled = scaler.transform(X_test_trans) # sklearn LinearRegression基准 model = LinearRegression() model.fit(X_train_trans,y_train) print("sklearn LinearRegression参数:") print(f"截距: {model.intercept_}") print(f"系数: {model.coef_}") # 手动梯度下降实现 n = len(X_train_scaled) β0 = 0.0 β1 = 0.0 β2 = 0.0 learning_rate = 0.01 num_iterations = 10000 for iteration in range(num_iterations): # 计算预测值(注意维度匹配,将y_prediction转为二维数组) y_prediction = β0 + β1 * X_train_scaled[:, 0].reshape(-1,1) + β2 * X_train_scaled[:, 1].reshape(-1,1) # 计算损失 loss = (1/(2*n)) * np.sum((y_train - y_prediction)**2) # 计算梯度 d_β0 = -(1 / n) * np.sum(y_train - y_prediction) d_β1 = -(1 / n) * np.sum((y_train - y_prediction) * X_train_scaled[:, 0].reshape(-1,1)) d_β2 = -(1 / n) * np.sum((y_train - y_prediction) * X_train_scaled[:, 1].reshape(-1,1)) # 更新参数 β0 -= learning_rate * d_β0 β1 -= learning_rate * d_β1 β2 -= learning_rate * d_β2 if iteration % 1000 == 0: print(f"Iteration {iteration}: Loss = {loss:.6f}") # 注意:由于特征做了标准化,需要将梯度下降得到的参数转换回原始特征尺度 # 还原系数:β1和β2需要除以scaler的标准差,截距需要调整 scaler_mean = scaler.mean_ scaler_scale = scaler.scale_ # 原始特征下的系数 β1_original = β1 / scaler_scale[0] β2_original = β2 / scaler_scale[1] β0_original = β0 - β1*(scaler_mean[0]/scaler_scale[0]) - β2*(scaler_mean[1]/scaler_scale[1]) print(f"\n还原后手动梯度下降参数(对应原始特征):") print(f"β0 = {β0_original[0]:.4f}, β1 = {β1_original[0]:.4f}, β2 = {β2_original[0]:.4f}")
结果验证
修改后的代码运行后,手动梯度下降得到的参数会和sklearn的LinearRegression结果高度接近,误差在可接受范围内。
内容的提问来源于stack exchange,提问作者Jit Roy
相关产品推荐
相关产品推荐

