基础线性回归无法拟合y=x+5模式,损失下降但拟合失效
线性回归模型拟合失败问题排查
问题描述
我基于Python实现了基础线性回归模型,目标拟合y=x+5的简单数据模式。目前观察到损失值持续下降,但模型始终无法准确拟合该模式;尝试了所有alpha取值,得到的最小损失为3.00。模型采用均方误差(mean squared error)作为损失函数。
代码实现
x=[] y=[] for i in range(1000): x.append(i) y.append(i+5) def compute_cost(x,y,w,b): m=len(x) total_cost=0 cost=0 for i in range(m): f_wb = w*x[i]+b cost += (f_wb - y[i])**2 total_cost = cost/(2*m) return total_cost def gradient_descent(x,y,iter=200): m=len(x) w=0 b=0 alpha=0.0000001 for i in range(iter): dj_dw,dj_db=Compute_gradient(x,y,w,b) w-=alpha*dj_dw b-=alpha*dj_db if(i%10==0): cost=compute_cost(x,y,w,b) print(f"cost at iteration {i} is {cost}") return w,b def Compute_gradient(x,y,w,b): m=len(x) dj_dw=0 dj_db=0 for i in range(m): y_predict=w*x[i]+b dj_dw +=(y_predict-y[i])*x[i] dj_db +=(y_predict-y[i]) dj_dw /=m dj_db /=m return dj_dw,dj_db w,b=gradient_descent(x,y) q=w*10+b print(q)
运行输出
cost at iteration 0 is 157869.16787934472 cost at iteration 10 is 80221.1618112745 cost at iteration 20 is 40765.111675992295 cost at iteration 30 is 20715.91825636264 cost at iteration 40 is 10528.123091503312 cost at iteration 50 is 5351.297861860537 cost at iteration 60 is 2720.746400189429 cost at iteration 70 is 1384.058239492339 cost at iteration 80 is 704.8336487635617 cost at iteration 90 is 359.6925315439946 cost at iteration 100 is 184.31255755759477 cost at iteration 110 is 95.19499412720666 cost at iteration 120 is 49.910803477032324 cost at iteration 130 is 26.900098542186733 cost at iteration 140 is 15.20744043888423 cost at iteration 150 is 9.265933534732453 cost at iteration 160 is 6.246816032905587 cost at iteration 170 is 4.712681193345336 cost at iteration 180 is 3.93312529712553 cost at iteration 190 is 3.537001070649068 10.064986594681566
问题根源分析
- 特征未做缩放处理:x的取值范围是0到999,数值量级远大于目标函数的截距项5。梯度下降中,w的梯度计算包含
x[i],导致dj_dw的量级远大于dj_db,单一学习率无法同时适配两个参数的更新速度——要么w更新太慢,要么b更新过度震荡。 - 迭代次数严重不足:仅200次迭代对于未缩放的数据来说,距离参数收敛到最优值还差得很远,只能让损失下降到3左右就进入缓慢收敛阶段。
- 学习率适配性差:不做特征缩放的情况下,调整alpha很难找到平衡值,小alpha收敛极慢,大alpha则会导致参数震荡无法收敛。
修复方案与优化代码
核心优化点
- 对x进行Z-score标准化,将特征转换为均值为0、方差为1的分布,让梯度下降的更新步长对所有参数更公平。
- 增加迭代次数,同时可以加入早停机制(当损失变化小于阈值时提前停止)。
- 标准化后可以使用更大的学习率,加快收敛速度。
修改后的代码
import numpy as np # 生成数据 x = np.arange(1000).tolist() y = [i + 5 for i in x] # 特征标准化 def normalize_features(x): x_np = np.array(x) mean = np.mean(x_np) std = np.std(x_np) return [(val - mean)/std for val in x], mean, std x_normalized, x_mean, x_std = normalize_features(x) def compute_cost(x,y,w,b): m = len(x) cost = 0 for i in range(m): f_wb = w * x[i] + b cost += (f_wb - y[i])**2 total_cost = cost / (2*m) return total_cost def Compute_gradient(x,y,w,b): m = len(x) dj_dw = 0 dj_db = 0 for i in range(m): y_predict = w * x[i] + b dj_dw += (y_predict - y[i]) * x[i] dj_db += (y_predict - y[i]) dj_dw /= m dj_db /= m return dj_dw, dj_db def gradient_descent(x,y,iter=1000, alpha=0.01): w = 0 b = 0 m = len(x) prev_cost = float('inf') for i in range(iter): dj_dw, dj_db = Compute_gradient(x,y,w,b) w -= alpha * dj_dw b -= alpha * dj_db cost = compute_cost(x,y,w,b) # 早停机制:损失变化小于1e-6时停止 if abs(prev_cost - cost) < 1e-6: print(f"Early stopping at iteration {i}, cost: {cost}") break prev_cost = cost if i % 100 == 0: print(f"cost at iteration {i} is {cost}") return w, b # 训练模型 w, b = gradient_descent(x_normalized, y) # 预测时需要对输入x做反标准化 def predict(x_val, w, b, x_mean, x_std): x_norm = (x_val - x_mean) / x_std return w * x_norm + b # 测试预测 q = predict(10, w, b, x_mean, x_std) print(f"预测x=10的结果:{q}") print(f"最优参数w: {w}, b: {b}")
优化后效果
标准化后,模型能快速收敛到接近最优的参数(对应原始x的w≈1、截距≈5),损失会降到接近0,预测值也会接近15(10+5)。
内容的提问来源于stack exchange,提问作者Krishan _13
相关产品推荐
相关产品推荐

