You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python从零实现线性回归 梯度下降不收敛求代码排查

梯度下降不收敛的核心问题排查

你的代码里有3个直接导致收敛失败的错误,逐个说明:

  • 损失函数实现完全错误
    你写的calc_naive_loss计算的是预测结果向量的L2范数,根本没有和真实标签y做差,完全无法衡量预测值和真实值的误差,打印出来的损失值没有任何参考意义。
    正确的平方误差损失(无正则项)应该为:预测值与真实值差值的平方和,即np.sum((y_pred - y)**2),如果用均方误差就再除以样本量。
  • 梯度未做缩放,步长与样本量冲突
    你推导的梯度公式X.T @ X @ w - X.T @ y是「残差平方和(未除以样本数、未乘1/2系数)」的梯度,这个梯度的量级和样本数成正比。你用的样本量是2010,梯度值会非常大,默认学习率0.001对应的步长过长,直接导致梯度爆炸,参数更新时来回震荡甚至越来越大,根本无法收敛。
  • 迭代逻辑存在漏洞
    你在类初始化时定义了最大迭代次数iterations参数,但梯度下降循环里完全没有用到这个终止条件,是纯死循环,一旦参数更新步长过大,权重变化量永远大于停止阈值,程序会无限运行直接跑飞。

修正方案

  1. 修复损失函数,加入真实标签的差值计算
  2. 梯度计算时除以样本量做缩放,适配现有学习率,也可以根据缩放后的梯度适当调大学习率加快收敛
  3. 在梯度下降循环中加入最大迭代次数的终止判断,避免死循环
  4. 修正了函数名的拼写错误(原代码中gradient_decent为拼写错误,正确应为gradient_descent),移除了冗余的入参

修正后的可运行代码

import numpy as np

class LinearReg:
    
    def __init__(self,with_reg=False,learning_rate=0.001,stopping_threshold=1e-6,iterations=100000):
        """
        线性回归模型初始化
        """
        self.with_reg=with_reg
        self.stopping_threshold=stopping_threshold
        self.iterations=iterations
        self.learning_rate=learning_rate
        
    def calc_naive_loss_gradient(self,weight_vector):
        """
        计算无正则项损失的梯度
        """
        y_pred = np.dot(self.x_updated, weight_vector)
        # 梯度除以样本量做缩放,避免梯度过大
        gradient = (2/self.n_samples) * np.dot(self.x_updated.transpose(), (y_pred - self.y))
        return gradient
    
    def calc_naive_loss(self,weight_vector):
        """
        计算均方误差损失
        """
        y_pred = np.dot(self.x_updated, weight_vector)
        return np.mean((y_pred - self.y)**2)
    
    def gradient_descent(self,weight_vector_new):
        """
        梯度下降迭代优化
        """
        for i in range(self.iterations):
            weight_vector_old = weight_vector_new.copy()            
            weight_vector_new = weight_vector_old - self.learning_rate * self.calc_naive_loss_gradient(weight_vector_old)
            
            # 按需打开迭代损失打印
            # if i % 1000 == 0:
            #     print(f'Iter {i}, Loss: {self.calc_naive_loss(weight_vector_new)}')

            dist_weights = np.sqrt(np.sum((weight_vector_new - weight_vector_old)**2))
            if dist_weights < self.stopping_threshold:
                break
        return weight_vector_new
    
    def fit(self,x,y):
        """
        训练线性回归模型
        """
        self.n_samples = x.shape[0]
        one_vector=np.ones(self.n_samples).reshape(self.n_samples,1)
        self.x_updated=np.concatenate((one_vector,x),axis=1)
        self.y=y
        
        # 初始化权重
        weight_vector=np.random.uniform(0,1,self.x_updated.shape[1])
        best_weight=self.gradient_descent(weight_vector_new=weight_vector.copy())
        
        print('Final loss: {}'.format(self.calc_naive_loss(weight_vector=best_weight)))
        self.w = best_weight

# 测试用例
from sklearn.datasets import make_regression
x, y = make_regression(n_features=5,n_samples=2010, noise=0, random_state=42)
a=LinearReg(learning_rate=0.1) # 梯度缩放后可使用更大学习率加快收敛
a.fit(x,y)

补充说明:如果使用的特征量级差异较大,需要先对特征做标准化处理,否则也会拖慢收敛速度。上述测试用例因为make_regression生成的特征本身服从标准正态分布,所以可以直接快速收敛。

内容的提问来源于stack exchange,提问作者Heisenberg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.01 02:01:34