You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于波士顿数据集的梯度下降算法实现维度不匹配报错求助

Fixing Dimension Mismatch in Multi-Feature Gradient Descent for Boston Dataset

Hey there, let's break down why you're hitting that dimension error and fix it up properly. The core issue is that your current code is built for single-feature linear regression, but you're feeding it 13 features (a 2D array). Let's adjust the gradient descent implementation to handle multi-variate linear regression correctly.

Why the Error Happens

When you use b0 + X*b1 with X as a (506,13) array and b1 as a single scalar, NumPy broadcasts this operation to create a (506,13) prediction array. But your target y is a 1D array of shape (506,), so subtracting the two (y - (b0 + X*b1)) fails because their shapes don't align. For multi-feature regression, we need a weight vector (one weight per feature) instead of a single b1.

Step-by-Step Fix

1. Adjust the Model for Multi-Feature Regression

The correct linear model for multiple features is:
$$\hat{y} = b_0 + w_1x_1 + w_2x_2 + ... + w_{13}x_{13}$$
We can rewrite this with matrix notation to simplify calculations:
$$\hat{y} = X_{augmented} \cdot w$$
Where:

  • $X_{augmented}$ is your original X with an extra column of 1s (to account for the bias term $b_0$)
  • $w$ is a weight vector of shape (14,) (13 feature weights + 1 bias term)

2. Rewrite the Gradient Descent Function

Here's the updated function that handles multi-feature data, plus we'll add feature scaling (critical for gradient descent to converge well with the Boston Dataset's varied feature scales):

import numpy as np
from sklearn.datasets import load_boston
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error
from sklearn.preprocessing import StandardScaler

# Load dataset
dataset = load_boston()
X = dataset.data
y = dataset.target

# Scale features (essential for stable gradient descent convergence)
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

# Add a column of 1s to X to include the bias term in our weight vector
X_augmented = np.hstack([np.ones((X_scaled.shape[0], 1)), X_scaled])

def gradient_descent(X, y, eta=0.01, n_iter=1000):
    '''Gradient descent implementation for multi-feature linear regression
    eta: Learning Rate
    n_iter: Number of iterations
    '''
    m = len(y)
    # Initialize weights: [b0, w1, w2, ..., w13]
    weights = np.zeros(X.shape[1])
    cost_list = []
    
    for i in range(n_iter):
        # Predict y using matrix multiplication (matches y's shape now)
        y_pred = X @ weights
        # Calculate error (valid since y and y_pred are both (506,))
        error = y - y_pred
        # Compute gradient using matrix operations (efficient for any feature count)
        gradient = (2/m) * X.T @ error
        # Update weights
        weights -= eta * gradient
        # Store current MSE
        cost = mean_squared_error(y, y_pred)
        cost_list.append(cost)
    
    return cost_list, weights

# Run gradient descent
cost_history, final_weights = gradient_descent(X_augmented, y)

# Plot cost reduction to verify convergence
plt.plot(range(len(cost_history)), cost_history)
plt.xlabel('Iterations')
plt.ylabel('Mean Squared Error')
plt.title('Cost Reduction During Gradient Descent')
plt.show()

# Compare with sklearn's LinearRegression for validation
lr = LinearRegression()
lr.fit(X_scaled, y)

print("Our final bias (b0):", final_weights[0])
print("Sklearn's intercept:", lr.intercept_)
print("\nOur feature weights:", final_weights[1:])
print("Sklearn's coefficients:", lr.coef_)

Key Changes Explained

  • Feature Scaling: We used StandardScaler to normalize features to mean 0 and variance 1. Without this, features with larger scales (like RM or LSTAT) would dominate gradient updates, making convergence slow or unstable.
  • Augmented X: Adding a column of 1s lets us include the bias term in the weight vector, simplifying matrix operations.
  • Weight Vector: Initialized weights as a zero vector matching the number of columns in the augmented X (14 elements: bias + 13 feature weights).
  • Matrix Gradient Calculation: Using X.T @ error computes the gradient efficiently for all weights at once, which works seamlessly for any number of features.

Why This Works

Now y_pred is a 1D array of shape (506,), matching y exactly—so the error calculation is valid. The gradient update uses linear algebra to compute the derivative of the MSE cost function for all weights simultaneously.

Bonus: Verify Convergence

The plotted cost history will show you if the MSE decreases steadily and plateaus, which means the algorithm is converging. If it oscillates or increases, tweak the learning rate (eta) or increase the number of iterations.

内容的提问来源于stack exchange,提问作者srbemr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:33:14