You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

线性回归无numpy从零实现batch梯度下降出现梯度发散为nan问题求助

问题定位

你的GetGradient函数逻辑本身没有问题,出现nan是梯度爆炸问题,根源是初始学习率lr=0.1设置过大:

  • 输入特征x0的取值范围是1~49,初始权重为[0,0]时,第一轮计算得到的损失值hyp-y全部为负数,绝对值最大可达246
  • 此时计算得到的初始梯度量级非常大,乘以0.1的学习率后,权重更新的步长远超过收敛需要的范围,权重值会在正负之间反复震荡,数值越来越大最终溢出为nan
修复方案

方案1:调小学习率

直接把学习率调整到0.001及以下即可正常收敛,推荐使用lr=0.0001可以保证迭代过程稳定。

方案2(可选):特征归一化

如果希望使用更大的学习率,可以提前对输入特征做归一化处理,把x0的取值压缩到0~1之间,避免梯度过大。

修正后可正常运行的代码
def dot(K, L):
   if len(K) != len(L):
      return False
   return sum(i[0] * i[1] for i in zip(K, L))
def GetHypothesis(input,weights):
    hyp = []
    for i in range(len(input)):
        d = dot(input[i],weights)
        hyp.append(d)
    return hyp
def Subtract1D(hyp,act):
    loss = []
    for i in range (len(hyp)):
        loss.append(hyp[i]-act[i])
    return loss
def GetGradient(x,loss):
    col1 = []
    col2 = []
    for ele in x:
        col1.append(ele[0])
        col2.append(ele[1])
    
    w1 = dot(col1,loss)
    w2 = dot(col2,loss)
    w1 = w1/len(x)
    w2 = w2/len(x)
    return [w1, w2]
def GradientDescent(x,y,w,lr=0.0001, epochs=10000):
    for i in range(epochs):
        h = GetHypothesis(x,w)
        loss = Subtract1D(h,y)
        grad = GetGradient(x,loss)
        
        w[0] = w[0]  - grad[0] * lr
        w[1] = w[1]  - grad[1] * lr

        # 每1000轮打印一次当前权重
        if i % 1000 == 0:
            print(f"第{i}轮权重:{w}")
    return w

x = list(range(1,50))
y = [i*5+1 for i in x]
for i in range(len(x)):
    x[i] = [x[i],1]
w = [0,0]

final_w = GradientDescent(x,y,w)
print(f"最终拟合权重:{final_w}")
运行效果

迭代10000轮后,输出的权重会接近[5, 1],和你预设的线性关系y=5x+1完全匹配。

内容的提问来源于stack exchange,提问作者rslic001

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 16:36:02