Python从零实现线性回归 梯度下降不收敛求代码排查
梯度下降不收敛的核心问题排查
你的代码里有3个直接导致收敛失败的错误,逐个说明:
- 损失函数实现完全错误
你写的calc_naive_loss计算的是预测结果向量的L2范数,根本没有和真实标签y做差,完全无法衡量预测值和真实值的误差,打印出来的损失值没有任何参考意义。
正确的平方误差损失(无正则项)应该为:预测值与真实值差值的平方和,即np.sum((y_pred - y)**2),如果用均方误差就再除以样本量。 - 梯度未做缩放,步长与样本量冲突
你推导的梯度公式X.T @ X @ w - X.T @ y是「残差平方和(未除以样本数、未乘1/2系数)」的梯度,这个梯度的量级和样本数成正比。你用的样本量是2010,梯度值会非常大,默认学习率0.001对应的步长过长,直接导致梯度爆炸,参数更新时来回震荡甚至越来越大,根本无法收敛。 - 迭代逻辑存在漏洞
你在类初始化时定义了最大迭代次数iterations参数,但梯度下降循环里完全没有用到这个终止条件,是纯死循环,一旦参数更新步长过大,权重变化量永远大于停止阈值,程序会无限运行直接跑飞。
修正方案
- 修复损失函数,加入真实标签的差值计算
- 梯度计算时除以样本量做缩放,适配现有学习率,也可以根据缩放后的梯度适当调大学习率加快收敛
- 在梯度下降循环中加入最大迭代次数的终止判断,避免死循环
- 修正了函数名的拼写错误(原代码中
gradient_decent为拼写错误,正确应为gradient_descent),移除了冗余的入参
修正后的可运行代码
import numpy as np class LinearReg: def __init__(self,with_reg=False,learning_rate=0.001,stopping_threshold=1e-6,iterations=100000): """ 线性回归模型初始化 """ self.with_reg=with_reg self.stopping_threshold=stopping_threshold self.iterations=iterations self.learning_rate=learning_rate def calc_naive_loss_gradient(self,weight_vector): """ 计算无正则项损失的梯度 """ y_pred = np.dot(self.x_updated, weight_vector) # 梯度除以样本量做缩放,避免梯度过大 gradient = (2/self.n_samples) * np.dot(self.x_updated.transpose(), (y_pred - self.y)) return gradient def calc_naive_loss(self,weight_vector): """ 计算均方误差损失 """ y_pred = np.dot(self.x_updated, weight_vector) return np.mean((y_pred - self.y)**2) def gradient_descent(self,weight_vector_new): """ 梯度下降迭代优化 """ for i in range(self.iterations): weight_vector_old = weight_vector_new.copy() weight_vector_new = weight_vector_old - self.learning_rate * self.calc_naive_loss_gradient(weight_vector_old) # 按需打开迭代损失打印 # if i % 1000 == 0: # print(f'Iter {i}, Loss: {self.calc_naive_loss(weight_vector_new)}') dist_weights = np.sqrt(np.sum((weight_vector_new - weight_vector_old)**2)) if dist_weights < self.stopping_threshold: break return weight_vector_new def fit(self,x,y): """ 训练线性回归模型 """ self.n_samples = x.shape[0] one_vector=np.ones(self.n_samples).reshape(self.n_samples,1) self.x_updated=np.concatenate((one_vector,x),axis=1) self.y=y # 初始化权重 weight_vector=np.random.uniform(0,1,self.x_updated.shape[1]) best_weight=self.gradient_descent(weight_vector_new=weight_vector.copy()) print('Final loss: {}'.format(self.calc_naive_loss(weight_vector=best_weight))) self.w = best_weight # 测试用例 from sklearn.datasets import make_regression x, y = make_regression(n_features=5,n_samples=2010, noise=0, random_state=42) a=LinearReg(learning_rate=0.1) # 梯度缩放后可使用更大学习率加快收敛 a.fit(x,y)
补充说明:如果使用的特征量级差异较大,需要先对特征做标准化处理,否则也会拖慢收敛速度。上述测试用例因为
make_regression生成的特征本身服从标准正态分布,所以可以直接快速收敛。
内容的提问来源于stack exchange,提问作者Heisenberg
相关产品推荐
相关产品推荐

