You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

逻辑回归Cost值不收敛求助:附代码及参数配置

Troubleshooting Your Logistic Regression Cost Convergence Issue

Hey there! Let’s dig into why your logistic regression cost isn’t converging. I’ve looked over your setup and spotted a few key issues that are likely throwing things off, plus some fixes to get you back on track.

Core Issues to Fix

1. Mixed-Up Data Dimensions

You mentioned 3 training samples (m=3) with 4 features each (n=4), but your X array is shaped (4,3)—that’s 4 rows and 3 columns. This reverses the sample-feature structure, which will break matrix multiplications and gradient calculations.

Logistic regression expects X to be (m, n): rows for samples, columns for features. So we need to transpose X to get (3,4). Also, your Y is (1,3)—it’s better to reshape it to (3,1) to match the sample count as column vectors, which makes gradient math cleaner.

2. Missing Sigmoid Activation

Logistic regression relies on the sigmoid function to map linear outputs to probabilities (0-1). I don’t see it in your code snippet—without it, you’re just doing linear regression, which won’t work for classification tasks and will lead to incorrect cost calculations.

3. Potential Cost Function Mismatch

If you’re using mean squared error (MSE) instead of cross-entropy loss, that’s another problem. MSE is non-convex for logistic regression, meaning gradient descent can get stuck in local minima instead of converging to the optimal solution. Cross-entropy is the standard loss function for this task.

4. Gradient Update Dimension Errors

With the original X shape, your weight updates (W) will have mismatched dimensions, so the gradients won’t adjust correctly. Fixing the data shape will resolve this, but you also need to make sure the gradient calculations use the right matrix transposes.

Corrected Code Example

Here’s a revised version of your code with all these fixes applied:

import numpy as np

# Fix data dimensions: 3 samples, 4 features each
X = np.array([[1,2,1],[1,1,0],[1,2,1],[1,0,2]]).T  # Now X.shape = (3,4)
Y = np.array([[0,1,0]]).T  # Y.shape = (3,1) to match sample count

# Parameter setup
m = X.shape[0]  # Number of samples: 3
n = X.shape[1]  # Number of features: 4
iterations = 100000
alpha = 0.05
b = 0  # Bias term
W = np.zeros((1, n))  # Weights: (1,4) to match feature count
cost_history = []

# Sigmoid activation function
def sigmoid(z):
    return 1 / (1 + np.exp(-z))

# Gradient descent loop
for i in range(iterations):
    # Calculate hypothesis (probabilities)
    z = np.dot(X, W.T) + b
    h = sigmoid(z)
    
    # Compute cross-entropy cost
    cost = -(1/m) * np.sum(Y * np.log(h) + (1 - Y) * np.log(1 - h))
    cost_history.append(cost)
    
    # Calculate gradients
    dW = (1/m) * np.dot((h - Y).T, X)
    db = (1/m) * np.sum(h - Y)
    
    # Update weights and bias
    W -= alpha * dW
    b -= alpha * db
    
    # Print progress every 10k iterations to check convergence
    if i % 10000 == 0:
        print(f"Iteration {i}, Cost: {cost:.4f}")

# Final results
print(f"\nFinal Cost: {cost_history[-1]:.4f}")
print(f"Final Weights: {W}")
print(f"Final Bias: {b}")

Extra Debugging Tips

  • Track Cost Over Iterations: Plotting cost_history will help you see if the cost is decreasing steadily (good) or oscillating/rising (bad). If it’s rising, try reducing the learning rate (alpha).
  • Feature Scaling: Even though your feature values are small (0-2), standardizing features (subtract mean, divide by standard deviation) can speed up convergence. Just make sure not to scale any constant bias term if you add one directly to X.
  • Check for Numerical Stability: If you get nan values, it might be because h is 0 or 1 (causing log(0) errors). You can add a small epsilon (like 1e-10) to h and 1-h to avoid this.

内容的提问来源于stack exchange,提问作者D. Wei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:48:17