You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python梯度下降线性回归计算异常求助:吴恩达课程学习遇阻

问题分析与修复方案

嘿,作为刚接触Python和机器学习的新手,你能写出这样的代码已经很棒了!咱们来一步步找出代码里的问题,让梯度下降正常工作起来。

核心错误:梯度计算逻辑错误

你的mean_error函数是问题的关键——当计算theta₁的梯度时,你把累加和每一步都乘以当前的x[i],这和正确的梯度公式完全不符。正确的theta₁梯度应该是对每个样本计算(h(xᵢ)-yᵢ)*xᵢ然后求和,而不是先把所有(h(xᵢ)-yᵢ)加起来再逐个乘x[i]。

看你的这段代码:

def mean_error(a, b, factor):
    sum_mean = 0
    for i in range(m):
        sum_mean += (theta_0 + theta_1 * a[i]) - b[i] # h(x) = (theta0 + theta1 * x) - y
        if factor:
            sum_mean *= a[i]
    return sum_mean

当factor=True时,每次循环都会把当前的累加和乘以x[i],这会导致结果变成类似((((h₀-y₀)*x₀) + (h₁-y₁))*x₁ + (h₂-y₂))*x₂...,完全偏离了我们需要的sum((h(xᵢ)-yᵢ)*xᵢ)。

修复方式

把factor的判断移到累加的步骤里,也就是每个项根据factor决定是否乘以x[i],然后再累加:

def mean_error(a, b, factor):
    sum_mean = 0
    for i in range(m):
        error = (theta_0 + theta_1 * a[i]) - b[i]
        if factor:
            sum_mean += error * a[i]
        else:
            sum_mean += error
    return sum_mean

其他需要注意的小问题

  • theta初始化的范围
    你用np.random.randint(low=2, high=5)初始化theta₀和theta₁,对于y=x这个数据集来说,正确的参数是theta₀=0,theta₁=1,初始值偏离有点大。虽然梯度下降最终能收敛,但可以改成更合理的初始化,比如用小的随机浮点数:
theta_1 = np.random.randn()  # 正态分布随机数
theta_0 = np.random.randn()
  • 绘图效率优化
    在迭代循环里每次ax.clear()然后重新绘制,虽然能运行,但可以提前把原始数据的散点图画好,只更新拟合线:
fig = plt.figure()
ax = fig.add_subplot(111)
# 先画一次原始数据
ax.plot(x, y, linestyle='None', marker='o')
# 初始化拟合线对象
line, = ax.plot(x, theta_0 + theta_1*x)
for i in range(iterations):
    theta_0, theta_1 = perform_cal(theta_0, theta_1, m)
    # 只更新拟合线的数据
    line.set_ydata(theta_0 + theta_1*x)
    fig.canvas.draw()
    fig.canvas.flush_events()  # 确保绘图实时更新
  • numpy向量化计算(可选)
    既然已经用了numpy,完全可以抛弃循环,用向量化操作来加速梯度计算,这也是吴恩达课程里会提到的技巧。比如把mean_error和perform_cal改成:
def perform_cal(theta_0, theta_1, m):
    h = theta_0 + theta_1 * x
    error = h - y
    gradient_0 = np.mean(error)
    gradient_1 = np.mean(error * x)
    temp_0 = theta_0 - learning_rate * gradient_0
    temp_1 = theta_1 - learning_rate * gradient_1
    return temp_0, temp_1

这样不仅代码更简洁,运行速度也更快,尤其当数据集大的时候。

修复后的完整代码

把上面的修改整合后,代码应该能正常收敛到theta₀≈0,theta₁≈1:

import numpy as np
import matplotlib.pyplot as plt
plt.ion()

x = [1,2,3,4,5]
y = [1,2,3,4,5]

def Gradient_Descent(x, y, learning_rate, iterations):
    # 更合理的参数初始化
    theta_1 = np.random.randn()
    theta_0 = np.random.randn()
    m = x.shape[0]

    def perform_cal(theta_0, theta_1, m):
        # 向量化计算梯度
        h = theta_0 + theta_1 * x
        error = h - y
        gradient_0 = np.mean(error)
        gradient_1 = np.mean(error * x)
        temp_0 = theta_0 - learning_rate * gradient_0
        temp_1 = theta_1 - learning_rate * gradient_1
        return temp_0 , temp_1

    fig = plt.figure()
    ax = fig.add_subplot(111)
    ax.plot(x, y, linestyle='None', marker='o')
    line, = ax.plot(x, theta_0 + theta_1*x)

    for i in range(iterations):
        theta_0, theta_1 = perform_cal(theta_0, theta_1, m)
        line.set_ydata(theta_0 + theta_1*x)
        fig.canvas.draw()
        fig.canvas.flush_events()
    
    print(f"最终参数:theta0={theta_0:.4f}, theta1={theta_1:.4f}")

x = np.array(x)
y = np.array(y)
Gradient_Descent(x,y, 0.1, 500)
input("Press enter to close program")

运行这段代码,你会看到拟合线逐渐靠近原始数据,最终收敛到正确的直线上。

内容的提问来源于stack exchange,提问作者Expressingx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:24:05