Python梯度下降线性回归计算异常求助:吴恩达课程学习遇阻
问题分析与修复方案
嘿,作为刚接触Python和机器学习的新手,你能写出这样的代码已经很棒了!咱们来一步步找出代码里的问题,让梯度下降正常工作起来。
核心错误:梯度计算逻辑错误
你的mean_error函数是问题的关键——当计算theta₁的梯度时,你把累加和每一步都乘以当前的x[i],这和正确的梯度公式完全不符。正确的theta₁梯度应该是对每个样本计算(h(xᵢ)-yᵢ)*xᵢ然后求和,而不是先把所有(h(xᵢ)-yᵢ)加起来再逐个乘x[i]。
看你的这段代码:
def mean_error(a, b, factor): sum_mean = 0 for i in range(m): sum_mean += (theta_0 + theta_1 * a[i]) - b[i] # h(x) = (theta0 + theta1 * x) - y if factor: sum_mean *= a[i] return sum_mean
当factor=True时,每次循环都会把当前的累加和乘以x[i],这会导致结果变成类似((((h₀-y₀)*x₀) + (h₁-y₁))*x₁ + (h₂-y₂))*x₂...,完全偏离了我们需要的sum((h(xᵢ)-yᵢ)*xᵢ)。
修复方式
把factor的判断移到累加的步骤里,也就是每个项根据factor决定是否乘以x[i],然后再累加:
def mean_error(a, b, factor): sum_mean = 0 for i in range(m): error = (theta_0 + theta_1 * a[i]) - b[i] if factor: sum_mean += error * a[i] else: sum_mean += error return sum_mean
其他需要注意的小问题
- theta初始化的范围
你用np.random.randint(low=2, high=5)初始化theta₀和theta₁,对于y=x这个数据集来说,正确的参数是theta₀=0,theta₁=1,初始值偏离有点大。虽然梯度下降最终能收敛,但可以改成更合理的初始化,比如用小的随机浮点数:
theta_1 = np.random.randn() # 正态分布随机数 theta_0 = np.random.randn()
- 绘图效率优化
在迭代循环里每次ax.clear()然后重新绘制,虽然能运行,但可以提前把原始数据的散点图画好,只更新拟合线:
fig = plt.figure() ax = fig.add_subplot(111) # 先画一次原始数据 ax.plot(x, y, linestyle='None', marker='o') # 初始化拟合线对象 line, = ax.plot(x, theta_0 + theta_1*x) for i in range(iterations): theta_0, theta_1 = perform_cal(theta_0, theta_1, m) # 只更新拟合线的数据 line.set_ydata(theta_0 + theta_1*x) fig.canvas.draw() fig.canvas.flush_events() # 确保绘图实时更新
- numpy向量化计算(可选)
既然已经用了numpy,完全可以抛弃循环,用向量化操作来加速梯度计算,这也是吴恩达课程里会提到的技巧。比如把mean_error和perform_cal改成:
def perform_cal(theta_0, theta_1, m): h = theta_0 + theta_1 * x error = h - y gradient_0 = np.mean(error) gradient_1 = np.mean(error * x) temp_0 = theta_0 - learning_rate * gradient_0 temp_1 = theta_1 - learning_rate * gradient_1 return temp_0, temp_1
这样不仅代码更简洁,运行速度也更快,尤其当数据集大的时候。
修复后的完整代码
把上面的修改整合后,代码应该能正常收敛到theta₀≈0,theta₁≈1:
import numpy as np import matplotlib.pyplot as plt plt.ion() x = [1,2,3,4,5] y = [1,2,3,4,5] def Gradient_Descent(x, y, learning_rate, iterations): # 更合理的参数初始化 theta_1 = np.random.randn() theta_0 = np.random.randn() m = x.shape[0] def perform_cal(theta_0, theta_1, m): # 向量化计算梯度 h = theta_0 + theta_1 * x error = h - y gradient_0 = np.mean(error) gradient_1 = np.mean(error * x) temp_0 = theta_0 - learning_rate * gradient_0 temp_1 = theta_1 - learning_rate * gradient_1 return temp_0 , temp_1 fig = plt.figure() ax = fig.add_subplot(111) ax.plot(x, y, linestyle='None', marker='o') line, = ax.plot(x, theta_0 + theta_1*x) for i in range(iterations): theta_0, theta_1 = perform_cal(theta_0, theta_1, m) line.set_ydata(theta_0 + theta_1*x) fig.canvas.draw() fig.canvas.flush_events() print(f"最终参数:theta0={theta_0:.4f}, theta1={theta_1:.4f}") x = np.array(x) y = np.array(y) Gradient_Descent(x,y, 0.1, 500) input("Press enter to close program")
运行这段代码,你会看到拟合线逐渐靠近原始数据,最终收敛到正确的直线上。
内容的提问来源于stack exchange,提问作者Expressingx
相关产品推荐
相关产品推荐

