基于SGD的Python神经网络在Moons数据集上不收敛问题求助
Hey there! Let's break down why your SGD isn't converging and those predictions are stuck around 0.9 (or all 1s) — this is a super common pitfall, so we'll work through the most likely culprits step by step.
Common Causes & Fixes
1. Mismatched Loss Function and Output Activation
Binary classification relies on pairing the right activation with the right loss:
- If you're using a
sigmoidoutput layer (which squashes values to 0-1 for binary labels), you must use binary cross-entropy loss (not MSE, for example). - Double-check your loss formula: it should be:
loss = -y * log(y_pred) - (1-y) * log(1-y_pred)
(Add a small epsilon like1e-10to avoidlog(0)errors!) - If your loss calculation is missing one of these terms, the gradient signal will be skewed, leading to lopsided predictions that never correct.
2. Poor Weight Initialization
If your weights are initialized to large positive values, the sigmoid activation will immediately hit its saturated region (output ~1), where gradients are nearly zero. SGD can't update weights when gradients vanish!
- Fix: Use small random initializations, like:
Xavier or He initialization also work great for keeping activations in a non-saturated range.W = np.random.normal(0, 0.01, (input_dim, output_dim)) b = np.zeros((1, output_dim))
3. Learning Rate Misconfiguration
You mentioned increasing the learning rate leads to all predictions being 1 — that's a classic sign of a learning rate that's too large. Oversized updates can cause weights to "jump" past the optimal solution, even pushing activations permanently into sigmoid's saturated end.
- Fix: Test log-scaled learning rates starting from
1e-4up to1e-2. Avoid jumping straight to large values; small, consistent updates work better for SGD.
4. Missing Data Preprocessing
SGD is extremely sensitive to feature scale. If your x1 and x2 features have different ranges, the gradient updates will prioritize the larger-scale feature, throwing off convergence.
- Fix: Normalize or standardize your features:
# Standardize to mean 0, std 1 X = (X - X.mean(axis=0)) / X.std(axis=0)
5. Incorrect Gradient Calculation
This is likely the issue you suspected — let's verify the backprop math:
For binary cross-entropy and sigmoid:
- Forward pass:
z = W @ X + b,y_pred = sigmoid(z) - Backward pass:
- Gradient of loss w.r.t. z:
dz = y_pred - y - Gradient w.r.t. W:
dW = (X.T @ dz) / num_samples(average over batch for stable updates) - Gradient w.r.t. b:
db = np.mean(dz)
- Gradient of loss w.r.t. z:
- If you have the sign wrong, forgot to average, or messed up matrix dimensions, your weights will update in the wrong direction (or not at all).
- Pro tip: Compare your manual gradients to those from an auto-diff tool (like PyTorch) for a small batch — this will quickly spot formula errors.
6. Wasting Epochs on a Stalled Model
If your model is stuck in a saturated activation state (gradients near zero), increasing epochs won't help. Fix the initialization/loss/gradient issues first, then adjust epoch count as needed.
Working Example Snippet
Here's a minimal, corrected implementation to reference:
import numpy as np from sklearn.datasets import make_moons from sklearn.model_selection import train_test_split # Generate and preprocess data X, y = make_moons(n_samples=1000, noise=0.1, random_state=42) X = (X - X.mean(axis=0)) / X.std(axis=0) # Standardize y = y.reshape(-1, 1) X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2) # Initialize weights properly input_dim = 2 output_dim = 1 W = np.random.normal(0, 0.01, (input_dim, output_dim)) b = np.zeros((1, output_dim)) def sigmoid(z): return 1 / (1 + np.exp(-z)) def binary_cross_entropy(y_pred, y): eps = 1e-10 return -np.mean(y * np.log(y_pred + eps) + (1 - y) * np.log(1 - y_pred + eps)) # Train with SGD learning_rate = 0.01 epochs = 1000 for epoch in range(epochs): # Forward pass z = X_train @ W + b y_pred = sigmoid(z) loss = binary_cross_entropy(y_pred, y_train) # Backward pass dz = y_pred - y_train dW = (X_train.T @ dz) / len(X_train) db = np.mean(dz) # Update weights W -= learning_rate * dW b -= learning_rate * db if epoch % 100 == 0: print(f"Epoch {epoch} | Loss: {loss:.4f}") # Evaluate y_pred_test = sigmoid(X_test @ W + b) accuracy = np.mean((y_pred_test > 0.5) == y_test) print(f"Test Accuracy: {accuracy:.4f}")
Start by checking your gradient calculation and weight initialization first — those are the most frequent offenders here. If you're still stuck, sharing a snippet of your code will help narrow things down further!
内容的提问来源于stack exchange,提问作者Swaroop

