如何解决Scipy fmin_tnc在逻辑回归实现中的ValueError?
Hey, let's work through this error together! That ValueError about matrix dimensions comes from two key issues in your code—let's break them down and fix them step by step.
1. Wrong Parameter Order for fmin_tnc
The biggest problem here is how you've structured your cost and grad functions relative to how opt.fmin_tnc expects to call them.
Scipy's fmin_tnc passes the optimization variable (your theta) as the first argument to the cost/gradient functions. But your current cost(x,y,theta) puts theta last. When you run opt.fmin_tnc(func=cost, x0=initial_theta, args=(x, y)), the function ends up calling cost(initial_theta, x, y)—swapping the variables entirely. That's why you get a dimension mismatch: it's trying to multiply initial_theta (shape (3,1)) with y (shape (99,1)), which doesn't work.
2. Sigmoid Function Dimension Inconsistency
Your sigmoid function initializes h_of_z as a 1D array (np.zeros([z.shape[0]])), which can break matrix operations later if your input z is 2D (like the result of x@theta, which is (99,1)). This can lead to unintended shape changes when you compute h_of_x - y in the gradient.
Fixed Code
Here's the corrected version of your code with these issues addressed:
import numpy as np from scipy import optimize as opt # Simplified Sigmoid (preserves input dimensions) def sigmoid(z): return 1 / (1 + np.exp(-z)) # Updated cost function: theta is first parameter def cost(theta, x, y): m = y.shape[0] h_of_x = sigmoid(np.matmul(x, theta)) # Use np.sum for better stability with array dimensions term1 = np.sum(-y * np.log(h_of_x) - (1 - y) * np.log(1 - h_of_x)) J = term1 / m return J # Updated gradient function: theta is first parameter def grad(theta, x, y): m = y.shape[0] h_of_x = sigmoid(x @ theta) grad = (x.T @ (h_of_x - y)) / m return grad # Add intercept term (assuming X is your original feature matrix, y is labels) x = np.hstack([np.ones_like(y), X[:, 0:2]]) [m, n] = x.shape initial_theta = np.zeros([n, 1]) # Call fmin_tnc with flattened theta (fmin_tnc prefers 1D arrays for x0) # Also pass the gradient function for faster convergence result = opt.fmin_tnc(func=cost, x0=initial_theta.flatten(), args=(x, y), fprime=grad)
Key Notes:
- We moved
thetato be the first argument in bothcostandgradto matchfmin_tnc's requirements. - The sigmoid function is simplified to preserve the input's shape (no need to pre-initialize an array).
- We use
initial_theta.flatten()to convert the 2D array to a 1D array—fmin_tncexpects this format for the initial guess, and it avoids potential warnings. - Adding
fprime=gradletsfmin_tncuse your custom gradient function instead of estimating it numerically, which is faster and more accurate.
内容的提问来源于stack exchange,提问作者Arun_Ramji_Shanmugam

