机器学习:计算代价函数误差时得到0尺寸矩阵的问题求助
Hey there, let's dig into why you're getting a size 0 matrix when calculating your cost function error. Looking at your code snippet, there are a few common pitfalls that could be causing this—let's break them down with practical fixes:
1. Inconsistent Shapes Between Arrays and Matrices
Numpy behaves unpredictably when mixing regular arrays and matrices. For example, if X is a numpy array (not a matrix), X[:,0] returns a 1D array instead of a column vector. Multiplying this with the column matrix result of sigmoid(X*theta)-y gives a matrix instead of the scalar you need for gradient updates. This shape mismatch can trickle down to your cost function and lead to empty matrices.
Fix: Convert X and y to matrices explicitly to ensure consistent column/row shapes:
def graD(X,y,alpha,s0,numda): X = np.asmatrix(X) # Convert X to matrix first y = np.asmatrix(y).T # Ensure y is a column matrix m = np.size(X,0) n = np.size(X,1) X0 = X[:,0] # Now this is a (m,1) column matrix X1 = X[:,1:] theta = np.asmatrix(np.zeros(n)).T # (n,1) matrix # Rest of your code...
2. Broken Theta Concatenation
Your code cuts off at theta=np.vstack((np.asmatrix(theta0),np.a...—I assume you meant to stack theta0 and theta1 back into a single theta matrix. If this line is incomplete (like using np.a instead of np.asmatrix(theta1)), it could result in an empty or malformed theta matrix, causing issues in cost function calculations.
Fix: Correct the concatenation line to properly stack the theta components:
theta = np.vstack((theta0, theta1))
Since theta0 is a (1,1) matrix and theta1 is a (n-1,1) matrix, this will give you the correct (n,1) theta matrix.
3. Missing Sum in Cost Function Calculation
If your cost function doesn't sum over all samples, you might end up with a matrix instead of a scalar. In edge cases with unexpected shapes, this can lead to a size 0 matrix. Here's a correct implementation of the regularized logistic regression cost function:
def cost_function(X, y, theta, numda): m = X.shape[0] h = sigmoid(X * theta) # Cross-entropy term (matrix multiplication gives scalar) cross_entropy = -y.T @ np.log(h) - (1 - y).T @ np.log(1 - h) # Regularization term (exclude theta0) reg_term = (numda / (2 * m)) * np.sum(np.square(theta[1:])) # Average cross-entropy plus regularization return (np.sum(cross_entropy) / m) + reg_term
This ensures you get a scalar cost value instead of a matrix.
Debugging Tip
To pinpoint shape issues early, print the shape of key variables at each step:
print(f"X shape: {X.shape}, y shape: {y.shape}, theta shape: {theta.shape}") print(f"sigmoid(X*theta) shape: {sigmoid(X*theta).shape}")
This will help you catch mismatches before they lead to empty matrices.
内容的提问来源于stack exchange,提问作者Zhihao ZHAO

