手动实现波士顿数据集线性回归SGD时参数不收敛问题求助
Hey there! Let's break down why your SGD implementation is blowing up weights instead of converging—there are several key mistakes we can fix step by step:
1. Gradients Never Reset Between Batches
You’re accumulating m_deriv and b_deriv across iterations without resetting them to 0 each time. This means gradients keep piling up, making your weight updates larger and larger until values explode.
Fix: Add m_deriv = 0 and b_deriv = 0 right after sampling your mini-batch to start fresh each iteration.
2. No Gradient Averaging for Mini-Batches
You’re using a mini-batch of 100 samples but not dividing the accumulated gradients by the batch size (N). This scales your gradients by the batch size, making updates way too aggressive (especially with your initial learning rate of 1).
Fix: After calculating the accumulated gradients, divide m_deriv and b_deriv by N before applying the learning rate.
3. Using Initial Weights Instead of Updated Ones
You initialize w0_random as your starting weights, but in the gradient calculation loop, you’re still using w0_random instead of the current updated weights (w0). This means you’re not actually learning—you’re computing gradients against the original random weights every time!
Fix: Replace np.dot(xi[i], w0_random) with np.dot(xi[i], w0) in both gradient calculation lines.
4. Unscaled Features
The Boston Housing dataset has features with wildly different scales (e.g., CRIM ranges from 0 to 88, while RM ranges from 3 to 9). Without standardizing these features, gradients will be dominated by large-scale features, making convergence nearly impossible.
Fix: Standardize your features (subtract mean, divide by standard deviation) before training using sklearn.preprocessing.StandardScaler.
5. Overly Aggressive Learning Rate
An initial learning rate of 1 is way too high for regression problems, and halving it every iteration will make the learning rate drop to near-zero extremely fast, preventing convergence before weights stabilize.
Fix: Start with a smaller initial learning rate (e.g., 0.01) and use a gradual decay (like multiplying by 0.95 each iteration) instead of halving.
6. Unrealistic Termination Condition
Checking if w0 == w1 exactly will almost never trigger due to floating-point precision errors. You’ll either get stuck in an infinite loop or your weights will blow up before this condition is met.
Fix: Check if the maximum absolute difference between w0 and w1 is below a small threshold (e.g., 1e-6), or add a maximum iteration count to prevent infinite loops.
Here’s your code with all fixes applied, plus some best practices:
import numpy as np import pandas as pd from sklearn.datasets import load_boston from sklearn.preprocessing import StandardScaler # Load and prepare data boston = load_boston() bos = pd.DataFrame(boston.data, columns=boston.feature_names) bos['price'] = boston.target # Standardize features to fix scale issues scaler = StandardScaler() scaled_features = scaler.fit_transform(bos.drop('price', axis=1)) bos_scaled = pd.DataFrame(scaled_features, columns=boston.feature_names) bos_scaled['price'] = bos['price'] # Initialize training parameters learning_rate = 0.01 max_iterations = 10000 # Prevent infinite loop tolerance = 1e-6 # Convergence threshold it = 0 # Initialize weights/bias with normal distribution (better for convergence) w0 = np.random.randn(13, 1) b0 = np.random.randn() while it < max_iterations: # Reset gradients for each new batch m_deriv = 0 b_deriv = 0 # Sample mini-batch df_sample = bos_scaled.sample(100) price = np.asmatrix(df_sample['price']).T # Fix shape to column vector xi = np.asmatrix(df_sample.drop('price', axis=1)) N = len(xi) # Calculate gradients over the mini-batch for i in range(N): y_pred = np.dot(xi[i], w0) + b0 error = price[i] - y_pred m_deriv += -2 * xi[i].T * error b_deriv += -2 * error # Average gradients by batch size m_deriv /= N b_deriv /= N # Update weights and bias w1 = w0 - learning_rate * m_deriv b1 = b0 - learning_rate * b_deriv # Check for convergence weight_diff = np.max(np.abs(w1 - w0)) bias_diff = np.abs(b1 - b0) if weight_diff < tolerance and bias_diff < tolerance: print(f"Converged after {it+1} iterations!") break # Update parameters for next iteration w0 = w1 b0 = b1 learning_rate *= 0.95 # Gradual learning rate decay it += 1 print("Final weights:\n", w0) print("Final bias:", b0)
- I switched to
np.random.randnfor initialization (normal distribution centered at 0) instead ofnp.random.rand—this helps with faster convergence since weights start closer to optimal values. - Fixed the shape of the
pricematrix to match predictions, avoiding broadcasting errors. - Added a maximum iteration count to prevent the code from running indefinitely if convergence is slow.
内容的提问来源于stack exchange,提问作者learner

