You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Logistic回归实现问题:np.log遇0报错与scipy.fmin_tnc形状不匹配

解决Andrew Ng逻辑回归实现中的两个问题

问题2:scipy.optimize.fmin_tnc的形状不匹配错误

这个问题的核心是fmin_tnc对输入参数的形状有明确要求:它期望初始参数x0是一维数组,但你定义的theta是(1,3)的矩阵。同时,代价函数和梯度函数接收的theta也需要匹配矩阵乘法的维度,否则就会触发形状不匹配的报错。

修改方案:

  1. 调整初始参数形状:调用fmin_tnc时,把x0改成theta.flatten(),将矩阵转成一维数组。
  2. 在函数内统一theta形状:在costfunction和gradient开头,把传入的一维theta转成(1,3)的矩阵,确保矩阵乘法维度正确;梯度函数最后要返回一维数组,符合fmin_tnc的要求。

修改后的代码片段:

def costfunction(theta,X,Y):
    # 把传入的一维theta转为(1,3)矩阵
    theta = np.matrix(theta).reshape((1,3))
    m = len(Y)
    hypothesis = sigmoid(np.dot(X, theta.T))
    # 顺便解决问题1的log(0)问题,加入clip限制取值范围
    hypothesis = np.clip(hypothesis, 1e-10, 1 - 1e-10)
    error = (-Y * np.log(hypothesis)) - ((1-Y) * np.log(1 - hypothesis))
    return np.sum(error) / m

def gradient(theta,X,Y):
    theta = np.matrix(theta).reshape((1,3))
    m = len(Y)
    error = sigmoid(X * theta.T) - Y
    # 用向量运算替代循环,更高效
    grad = (X.T @ error) / m
    # 返回一维数组
    return grad.flatten()

# 调用fmin_tnc时传入扁平化的theta
import scipy.optimize as opt
result = opt.fmin_tnc(func=costfunction, x0=theta.flatten(), fprime=gradient, args=(X,Y))

这样就能解决形状不匹配的错误,调用costfunction(result[0], X, Y)也能正常计算代价。


问题1:alpha=0.01时的数值错误与alpha=0.001收敛过慢

错误原因:

  • 取对数遇到零值:当alpha较大时,theta更新幅度过大,导致sigmoid(X*theta.T)的输出趋近于1(或0),此时1 - hypothesis(或hypothesis)会变成0,触发log(0)的错误。
  • 收敛过慢:你的特征X没有做标准化,不同特征的数值范围差异可能很大,导致梯度下降的步长难以调整——alpha小了收敛慢,alpha大了又会出现数值饱和甚至震荡。

解决步骤:

  1. 避免log(0)的数值问题:在计算完hypothesis后,用np.clip把值限制在[1e-10, 1-1e-10]之间,这样取对数时就不会遇到0或1。
  2. 特征标准化:对X的非偏置项(第1、2列)做标准化处理(减去均值除以标准差),让所有特征的数值范围接近,这样就能使用更大的alpha,大幅减少迭代次数。

修改后的梯度下降相关代码:

# 先对X做特征标准化(跳过第一列的ones)
X[:,1] = (X[:,1] - np.mean(X[:,1])) / np.std(X[:,1])
X[:,2] = (X[:,2] - np.mean(X[:,2])) / np.std(X[:,2])

def sigmoid(z): 
    return 1/(1+np.exp(-z))

def costfunction(theta,X,Y):
    theta = np.matrix(theta).reshape((1,3))
    m = len(Y)
    hypothesis = sigmoid(np.dot(X, theta.T))
    hypothesis = np.clip(hypothesis, 1e-10, 1 - 1e-10)
    error = (-Y * np.log(hypothesis)) - ((1-Y) * np.log(1 - hypothesis))
    return np.sum(error) / m

def gradientdescent(X,Y,theta,alpha,iters):
    parameters = 3
    temp = np.matrix(np.zeros(theta.shape))
    cost = np.zeros(iters)
    m = len(Y)
    for i in range(iters):
        error = sigmoid(X*theta.T) - Y
        # 向量运算替代循环,提升效率
        temp = theta - (alpha/m) * (error.T @ X)
        theta = temp
        cost[i] = costfunction(theta,X,Y)
    return theta, cost

# 现在可以用更大的alpha,比如0.1,迭代1000次就足够
alpha = 0.1
iters = 1000
param,cost = gradientdescent(X,Y,theta,alpha,iters)

这样修改后,alpha=0.1时不会出现数值错误,而且迭代1000次就能让代价从0.693快速降到0.2左右,不需要百万次迭代。


内容的提问来源于stack exchange,提问作者Anwaar Khalid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 19:12:54