Logistic回归实现问题:np.log遇0报错与scipy.fmin_tnc形状不匹配
解决Andrew Ng逻辑回归实现中的两个问题
问题2:scipy.optimize.fmin_tnc的形状不匹配错误
这个问题的核心是fmin_tnc对输入参数的形状有明确要求:它期望初始参数x0是一维数组,但你定义的theta是(1,3)的矩阵。同时,代价函数和梯度函数接收的theta也需要匹配矩阵乘法的维度,否则就会触发形状不匹配的报错。
修改方案:
- 调整初始参数形状:调用
fmin_tnc时,把x0改成theta.flatten(),将矩阵转成一维数组。 - 在函数内统一theta形状:在
costfunction和gradient开头,把传入的一维theta转成(1,3)的矩阵,确保矩阵乘法维度正确;梯度函数最后要返回一维数组,符合fmin_tnc的要求。
修改后的代码片段:
def costfunction(theta,X,Y): # 把传入的一维theta转为(1,3)矩阵 theta = np.matrix(theta).reshape((1,3)) m = len(Y) hypothesis = sigmoid(np.dot(X, theta.T)) # 顺便解决问题1的log(0)问题,加入clip限制取值范围 hypothesis = np.clip(hypothesis, 1e-10, 1 - 1e-10) error = (-Y * np.log(hypothesis)) - ((1-Y) * np.log(1 - hypothesis)) return np.sum(error) / m def gradient(theta,X,Y): theta = np.matrix(theta).reshape((1,3)) m = len(Y) error = sigmoid(X * theta.T) - Y # 用向量运算替代循环,更高效 grad = (X.T @ error) / m # 返回一维数组 return grad.flatten() # 调用fmin_tnc时传入扁平化的theta import scipy.optimize as opt result = opt.fmin_tnc(func=costfunction, x0=theta.flatten(), fprime=gradient, args=(X,Y))
这样就能解决形状不匹配的错误,调用costfunction(result[0], X, Y)也能正常计算代价。
问题1:alpha=0.01时的数值错误与alpha=0.001收敛过慢
错误原因:
- 取对数遇到零值:当alpha较大时,theta更新幅度过大,导致
sigmoid(X*theta.T)的输出趋近于1(或0),此时1 - hypothesis(或hypothesis)会变成0,触发log(0)的错误。 - 收敛过慢:你的特征X没有做标准化,不同特征的数值范围差异可能很大,导致梯度下降的步长难以调整——alpha小了收敛慢,alpha大了又会出现数值饱和甚至震荡。
解决步骤:
- 避免log(0)的数值问题:在计算完hypothesis后,用
np.clip把值限制在[1e-10, 1-1e-10]之间,这样取对数时就不会遇到0或1。 - 特征标准化:对X的非偏置项(第1、2列)做标准化处理(减去均值除以标准差),让所有特征的数值范围接近,这样就能使用更大的alpha,大幅减少迭代次数。
修改后的梯度下降相关代码:
# 先对X做特征标准化(跳过第一列的ones) X[:,1] = (X[:,1] - np.mean(X[:,1])) / np.std(X[:,1]) X[:,2] = (X[:,2] - np.mean(X[:,2])) / np.std(X[:,2]) def sigmoid(z): return 1/(1+np.exp(-z)) def costfunction(theta,X,Y): theta = np.matrix(theta).reshape((1,3)) m = len(Y) hypothesis = sigmoid(np.dot(X, theta.T)) hypothesis = np.clip(hypothesis, 1e-10, 1 - 1e-10) error = (-Y * np.log(hypothesis)) - ((1-Y) * np.log(1 - hypothesis)) return np.sum(error) / m def gradientdescent(X,Y,theta,alpha,iters): parameters = 3 temp = np.matrix(np.zeros(theta.shape)) cost = np.zeros(iters) m = len(Y) for i in range(iters): error = sigmoid(X*theta.T) - Y # 向量运算替代循环,提升效率 temp = theta - (alpha/m) * (error.T @ X) theta = temp cost[i] = costfunction(theta,X,Y) return theta, cost # 现在可以用更大的alpha,比如0.1,迭代1000次就足够 alpha = 0.1 iters = 1000 param,cost = gradientdescent(X,Y,theta,alpha,iters)
这样修改后,alpha=0.1时不会出现数值错误,而且迭代1000次就能让代价从0.693快速降到0.2左右,不需要百万次迭代。
内容的提问来源于stack exchange,提问作者Anwaar Khalid
相关产品推荐
相关产品推荐

