逻辑回归Cost值不收敛求助:附代码及参数配置
Hey there! Let’s dig into why your logistic regression cost isn’t converging. I’ve looked over your setup and spotted a few key issues that are likely throwing things off, plus some fixes to get you back on track.
Core Issues to Fix
1. Mixed-Up Data Dimensions
You mentioned 3 training samples (m=3) with 4 features each (n=4), but your X array is shaped (4,3)—that’s 4 rows and 3 columns. This reverses the sample-feature structure, which will break matrix multiplications and gradient calculations.
Logistic regression expects X to be (m, n): rows for samples, columns for features. So we need to transpose X to get (3,4). Also, your Y is (1,3)—it’s better to reshape it to (3,1) to match the sample count as column vectors, which makes gradient math cleaner.
2. Missing Sigmoid Activation
Logistic regression relies on the sigmoid function to map linear outputs to probabilities (0-1). I don’t see it in your code snippet—without it, you’re just doing linear regression, which won’t work for classification tasks and will lead to incorrect cost calculations.
3. Potential Cost Function Mismatch
If you’re using mean squared error (MSE) instead of cross-entropy loss, that’s another problem. MSE is non-convex for logistic regression, meaning gradient descent can get stuck in local minima instead of converging to the optimal solution. Cross-entropy is the standard loss function for this task.
4. Gradient Update Dimension Errors
With the original X shape, your weight updates (W) will have mismatched dimensions, so the gradients won’t adjust correctly. Fixing the data shape will resolve this, but you also need to make sure the gradient calculations use the right matrix transposes.
Corrected Code Example
Here’s a revised version of your code with all these fixes applied:
import numpy as np # Fix data dimensions: 3 samples, 4 features each X = np.array([[1,2,1],[1,1,0],[1,2,1],[1,0,2]]).T # Now X.shape = (3,4) Y = np.array([[0,1,0]]).T # Y.shape = (3,1) to match sample count # Parameter setup m = X.shape[0] # Number of samples: 3 n = X.shape[1] # Number of features: 4 iterations = 100000 alpha = 0.05 b = 0 # Bias term W = np.zeros((1, n)) # Weights: (1,4) to match feature count cost_history = [] # Sigmoid activation function def sigmoid(z): return 1 / (1 + np.exp(-z)) # Gradient descent loop for i in range(iterations): # Calculate hypothesis (probabilities) z = np.dot(X, W.T) + b h = sigmoid(z) # Compute cross-entropy cost cost = -(1/m) * np.sum(Y * np.log(h) + (1 - Y) * np.log(1 - h)) cost_history.append(cost) # Calculate gradients dW = (1/m) * np.dot((h - Y).T, X) db = (1/m) * np.sum(h - Y) # Update weights and bias W -= alpha * dW b -= alpha * db # Print progress every 10k iterations to check convergence if i % 10000 == 0: print(f"Iteration {i}, Cost: {cost:.4f}") # Final results print(f"\nFinal Cost: {cost_history[-1]:.4f}") print(f"Final Weights: {W}") print(f"Final Bias: {b}")
Extra Debugging Tips
- Track Cost Over Iterations: Plotting
cost_historywill help you see if the cost is decreasing steadily (good) or oscillating/rising (bad). If it’s rising, try reducing the learning rate (alpha). - Feature Scaling: Even though your feature values are small (0-2), standardizing features (subtract mean, divide by standard deviation) can speed up convergence. Just make sure not to scale any constant bias term if you add one directly to
X. - Check for Numerical Stability: If you get
nanvalues, it might be becausehis 0 or 1 (causinglog(0)errors). You can add a small epsilon (like1e-10) tohand1-hto avoid this.
内容的提问来源于stack exchange,提问作者D. Wei

