You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

手动实现波士顿数据集线性回归SGD时参数不收敛问题求助

Hey there! Let's break down why your SGD implementation is blowing up weights instead of converging—there are several key mistakes we can fix step by step:

Key Issues in Your Code

1. Gradients Never Reset Between Batches

You’re accumulating m_deriv and b_deriv across iterations without resetting them to 0 each time. This means gradients keep piling up, making your weight updates larger and larger until values explode.

Fix: Add m_deriv = 0 and b_deriv = 0 right after sampling your mini-batch to start fresh each iteration.

2. No Gradient Averaging for Mini-Batches

You’re using a mini-batch of 100 samples but not dividing the accumulated gradients by the batch size (N). This scales your gradients by the batch size, making updates way too aggressive (especially with your initial learning rate of 1).

Fix: After calculating the accumulated gradients, divide m_deriv and b_deriv by N before applying the learning rate.

3. Using Initial Weights Instead of Updated Ones

You initialize w0_random as your starting weights, but in the gradient calculation loop, you’re still using w0_random instead of the current updated weights (w0). This means you’re not actually learning—you’re computing gradients against the original random weights every time!

Fix: Replace np.dot(xi[i], w0_random) with np.dot(xi[i], w0) in both gradient calculation lines.

4. Unscaled Features

The Boston Housing dataset has features with wildly different scales (e.g., CRIM ranges from 0 to 88, while RM ranges from 3 to 9). Without standardizing these features, gradients will be dominated by large-scale features, making convergence nearly impossible.

Fix: Standardize your features (subtract mean, divide by standard deviation) before training using sklearn.preprocessing.StandardScaler.

5. Overly Aggressive Learning Rate

An initial learning rate of 1 is way too high for regression problems, and halving it every iteration will make the learning rate drop to near-zero extremely fast, preventing convergence before weights stabilize.

Fix: Start with a smaller initial learning rate (e.g., 0.01) and use a gradual decay (like multiplying by 0.95 each iteration) instead of halving.

6. Unrealistic Termination Condition

Checking if w0 == w1 exactly will almost never trigger due to floating-point precision errors. You’ll either get stuck in an infinite loop or your weights will blow up before this condition is met.

Fix: Check if the maximum absolute difference between w0 and w1 is below a small threshold (e.g., 1e-6), or add a maximum iteration count to prevent infinite loops.

Corrected Code Example

Here’s your code with all fixes applied, plus some best practices:

import numpy as np
import pandas as pd
from sklearn.datasets import load_boston
from sklearn.preprocessing import StandardScaler

# Load and prepare data
boston = load_boston()
bos = pd.DataFrame(boston.data, columns=boston.feature_names)
bos['price'] = boston.target

# Standardize features to fix scale issues
scaler = StandardScaler()
scaled_features = scaler.fit_transform(bos.drop('price', axis=1))
bos_scaled = pd.DataFrame(scaled_features, columns=boston.feature_names)
bos_scaled['price'] = bos['price']

# Initialize training parameters
learning_rate = 0.01
max_iterations = 10000  # Prevent infinite loop
tolerance = 1e-6  # Convergence threshold
it = 0

# Initialize weights/bias with normal distribution (better for convergence)
w0 = np.random.randn(13, 1)
b0 = np.random.randn()

while it < max_iterations:
    # Reset gradients for each new batch
    m_deriv = 0
    b_deriv = 0
    
    # Sample mini-batch
    df_sample = bos_scaled.sample(100)
    price = np.asmatrix(df_sample['price']).T  # Fix shape to column vector
    xi = np.asmatrix(df_sample.drop('price', axis=1))
    N = len(xi)
    
    # Calculate gradients over the mini-batch
    for i in range(N):
        y_pred = np.dot(xi[i], w0) + b0
        error = price[i] - y_pred
        
        m_deriv += -2 * xi[i].T * error
        b_deriv += -2 * error
    
    # Average gradients by batch size
    m_deriv /= N
    b_deriv /= N
    
    # Update weights and bias
    w1 = w0 - learning_rate * m_deriv
    b1 = b0 - learning_rate * b_deriv
    
    # Check for convergence
    weight_diff = np.max(np.abs(w1 - w0))
    bias_diff = np.abs(b1 - b0)
    
    if weight_diff < tolerance and bias_diff < tolerance:
        print(f"Converged after {it+1} iterations!")
        break
    
    # Update parameters for next iteration
    w0 = w1
    b0 = b1
    learning_rate *= 0.95  # Gradual learning rate decay
    it += 1

print("Final weights:\n", w0)
print("Final bias:", b0)
Extra Notes
  • I switched to np.random.randn for initialization (normal distribution centered at 0) instead of np.random.rand—this helps with faster convergence since weights start closer to optimal values.
  • Fixed the shape of the price matrix to match predictions, avoiding broadcasting errors.
  • Added a maximum iteration count to prevent the code from running indefinitely if convergence is slow.

内容的提问来源于stack exchange,提问作者learner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:42:08