You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

手动实现梯度下降算法与Sklearn结果不符,请求排查错误

Hey there, let's dig into why your manual gradient descent implementation isn't matching Sklearn's SGDRegressor results. I've spotted several key issues in your code—let's break them down one by one:

1. Feature Scaling Mismatch

The biggest culprit here is missing feature normalization in your manual code. Your input feature X ranges from 0 to 9, while the bias term's feature (the column of ones) is always 1. This huge scale imbalance means the gradient updates for the intercept (theta_0) and coefficient (theta_1) happen at drastically different rates, preventing your model from converging to the optimal values as efficiently as Sklearn's implementation.

To fix this, standardize your X feature before running gradient descent:

# Standardize X to mean 0, standard deviation 1
X_scaled = (X - np.mean(X)) / np.std(X)
X_new = np.concatenate((X0, X_scaled), axis=1)

And match this in your Sklearn code by enabling normalization:

model = SGDRegressor(loss='squared_loss', learning_rate='constant', eta0=0.01, max_iter=1000, normalize=True)

2. Batch vs. Stochastic Gradient Descent

Your manual code uses Batch Gradient Descent (BGD)—it computes the gradient using all training samples every iteration. But Sklearn's SGDRegressor defaults to Stochastic Gradient Descent (SGD), which uses one random sample per iteration. This difference in optimization strategy means the two models take different paths to convergence, even with identical learning rates and iteration counts.

To align them, you can either:

  • Modify your manual code to use SGD (sample one data point per iteration for gradient calculation)
  • Force SGDRegressor to use BGD by setting batch_size to the total number of samples:
    model = SGDRegressor(loss='squared_loss', learning_rate='constant', eta0=0.01, max_iter=1000, normalize=True, batch_size=len(Data))
    

3. Incorrect Cost Function Calculation

Your cost function J is computed wrong—it stores per-sample losses instead of the total average cost:

J = loss/(2*m)  # This creates a vector of individual sample losses, not the total cost

The correct half-mean-squared-error cost function (used for easier differentiation) should be:

J = (1/(2*m)) * np.sum(loss ** 2)  # Sum squared losses, then scale by 1/(2m)

Your current code leads you to select a suboptimal theta set when you find the "minimum J"—you're picking the iteration where the first sample's loss was smallest, not where overall cost was minimized.

4. Initial Theta Setup

You're initializing theta with np.random.randint(low=0, high=1), which generates all zeros (since randint is left-closed, right-open—high=1 only allows 0). Sklearn's SGDRegressor initializes weights with small random values by default. To match, use:

theta = np.random.randn(X_new.shape[1]).reshape(-1, 1) * 0.01  # Small random initialization

5. Iteration Count Discrepancy

Your loop runs for iter in range(1, iterations)—this executes 999 iterations, but SGDRegressor uses max_iter=1000. Align the counts to ensure both models get the same number of updates:

for iter in range(iterations):  # Runs full 1000 iterations

Fixed Manual Implementation

Here's your corrected code with all fixes applied:

import pandas as pd
import numpy as np
import seaborn as sb
import matplotlib.pyplot as plt
from sklearn.linear_model import SGDRegressor

Data = pd.DataFrame({'X': list(np.arange(0,10,1)), 'Y': [1,3,2,5,7,8,8,9,10,12]})
sb.scatterplot(x='X', y='Y', data=Data)
plt.show()

# Generate bias term column
X0 = np.ones(len(Data)).reshape(-1,1)
X = Data.drop(['Y'], axis=1).values
# Standardize feature X
X_scaled = (X - np.mean(X)) / np.std(X)
X_new = np.concatenate((X0, X_scaled), axis=1)
Y = Data.loc[:,['Y']].values

# Initialize theta with small random values
theta = np.random.randn(X_new.shape[1]).reshape(-1,1) * 0.01

J_history = []
theta_history = [list(theta.flatten())]

# Gradient descent hyperparameters
iterations = 1000
alpha = 0.01
m = len(Y)

for iter in range(iterations):
    H = X_new.dot(theta)
    loss = H - Y
    # Correct cost function calculation
    J = (1/(2*m)) * np.sum(loss ** 2)
    J_history.append(J)
    # Compute gradient
    G = X_new.T.dot(loss) / m
    theta_new = theta - alpha * G
    theta_history.append(list(theta_new.flatten()))
    theta = theta_new

# Find iteration with minimum total cost
idx = J_history.index(min(J_history))
values = theta_history[idx+1]  # theta_history includes initial theta + all updates
print('intercept: {}, coeff: {}'.format(values[0], values[1]))

Corrected Sklearn Code (Matching BGD)

import pandas as pd
import numpy as np
import seaborn as sb
import matplotlib.pyplot as plt
from sklearn.linear_model import SGDRegressor

Data = pd.DataFrame({'X': list(np.arange(0,10,1)), 'Y': [1,3,2,5,7,8,8,9,10,12]})
sb.scatterplot(x='X', y='Y', data=Data)
plt.show()

# Use Batch Gradient Descent, normalize features, fix random state for consistency
model = SGDRegressor(loss='squared_loss', learning_rate='constant', eta0=0.01, max_iter=1000, normalize=True, batch_size=len(Data), random_state=42)
model.fit(Data['X'].values.reshape(-1,1), Data['Y'].values)
print('coeff: {}, intercept: {}'.format(model.coef_, model.intercept_))

After these changes, you'll see the coefficients and intercept from both implementations align closely!

内容的提问来源于stack exchange,提问作者Pavan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 08:37:46