梯度下降Python代码调试求助:输出结果与预期不符
梯度下降代码错误排查求助
我正在为课程编写梯度下降的Python代码,使用Google Colab完成编程项目,但代码运行结果与预期不符,怀疑问题出在求和部分,请求帮助排查错误。
代码
import numpy as np #Parameters lamb= 0.1 beta = 1 alpha = 1 #Label at z is +1 and -z is -1 #Optimal w = z / ||z||^2 and b = 0 z = np.array([1,2]) #Data Matrix and Labels x = np.vstack((z,-z)) y = np.array([1,-1]) #Random Initialization w = np.array([-0.34637426, 1.16720011]) b = np.array([-0.24424188]) m=len(z) sumw=0 sumb=0 #Insert your code here to compute the gradients and loss. You can #Use as many additional lines of code as needed (i.e., don't try too #hard to put the whole computation in one line) for i in range(1000): for j in range(m): sumb+=(y[j])/(1+np.exp((-beta)*(1-y[j]*(x[j]@w-b)))) sumw+=(y[j]*x[j])/(1+np.exp((-beta)*(1-y[j]*(x[j]@w-b)))) grad_w = 2 * lamb * w - (1/2)*sumw grad_b=(1/m)*sumb b-=alpha*grad_b w-=alpha*grad_w if i % 100 == 0: print(w,' ', b)
当前输出
[-0.00311842 1.48172206] [-0.19604564] [2.52108931 5.04217862] [5.53009896] [2.7530958 5.50619161] [3.85357184] [2.76527666 5.53055331] [0.82437199] [2.76756568 5.53513137] [-2.25691088] [2.79916861 5.59833721] [-5.09974894] [2.93920254 5.87840507] [-6.08686155] [3.00843697 6.01687395] [-4.78457688] [3.01993034 6.03986068] [-2.86787546] [3.02152889 6.04305778] [-0.86030793]
预期输出
[-0.34637426 1.16720011] [-0.24424188] [0.5989022 1.19780439] [-3.17437461e-06] [0.5989022 1.19780439] [-4.60193897e-11] [0.5989022 1.19780439] [-5.82867088e-16] [0.5989022 1.19780439] [-1.94289029e-16] [0.5989022 1.19780439] [-1.94289029e-16] [0.5989022 1.19780439] [-1.94289029e-16] [0.5989022 1.19780439] [-1.94289029e-16] [0.5989022 1.19780439] [-1.94289029e-16] [0.5989022 1.19780439] [-1.94289029e-16]
问题背景(对应原Problem 4 iii)
该问题是带L2正则化的逻辑回归变种,目标是通过梯度下降优化参数w和b,最优解为w = z / ||z||²,b = 0。
错误排查与修复
- 求和变量未重置:
sumw和sumb在循环外初始化后,每次迭代没有重置为0,导致梯度计算持续累加历史求和结果,引发参数发散。需将sumw=0和sumb=0移到外层循环内部,每次迭代前重置。 - 梯度符号错误:
grad_b的计算缺少负号,根据损失函数梯度推导,正确梯度应为-(1/m)*sumb,否则参数更新方向错误,无法收敛到最优解。 - 预测项符号错误:逻辑回归预测项应为
x[j]@w + b,原代码写成x[j]@w - b,导致模型预测逻辑错误,影响梯度计算正确性。
修复后的代码:
import numpy as np #Parameters lamb= 0.1 beta = 1 alpha = 1 #Label at z is +1 and -z is -1 #Optimal w = z / ||z||^2 and b = 0 z = np.array([1,2]) #Data Matrix and Labels x = np.vstack((z,-z)) y = np.array([1,-1]) #Random Initialization w = np.array([-0.34637426, 1.16720011]) b = np.array([-0.24424188]) m = len(y) # 明确样本数 #Insert your code here to compute the gradients and loss. You can #Use as many additional lines of code as needed (i.e., don't try too #hard to put the whole computation in one line) for i in range(1000): sumw = 0 # 每次迭代前重置求和变量 sumb = 0 for j in range(m): # 修正预测项符号 exp_term = np.exp((-beta)*(1 - y[j]*(x[j]@w + b))) sumb += y[j] / (1 + exp_term) sumw += y[j] * x[j] / (1 + exp_term) grad_w = 2 * lamb * w - (1/m)*sumw grad_b = -(1/m)*sumb # 修正梯度符号 b -= alpha*grad_b w -= alpha*grad_w if i % 100 == 0: print(w,' ', b)
运行修复后的代码,参数会快速收敛到预期的最优值。
内容的提问来源于stack exchange,提问作者Cole Hendrickson
相关产品推荐
相关产品推荐

