基于波士顿数据集的梯度下降算法实现维度不匹配报错求助
Hey there, let's break down why you're hitting that dimension error and fix it up properly. The core issue is that your current code is built for single-feature linear regression, but you're feeding it 13 features (a 2D array). Let's adjust the gradient descent implementation to handle multi-variate linear regression correctly.
Why the Error Happens
When you use b0 + X*b1 with X as a (506,13) array and b1 as a single scalar, NumPy broadcasts this operation to create a (506,13) prediction array. But your target y is a 1D array of shape (506,), so subtracting the two (y - (b0 + X*b1)) fails because their shapes don't align. For multi-feature regression, we need a weight vector (one weight per feature) instead of a single b1.
Step-by-Step Fix
1. Adjust the Model for Multi-Feature Regression
The correct linear model for multiple features is:
$$\hat{y} = b_0 + w_1x_1 + w_2x_2 + ... + w_{13}x_{13}$$
We can rewrite this with matrix notation to simplify calculations:
$$\hat{y} = X_{augmented} \cdot w$$
Where:
- $X_{augmented}$ is your original
Xwith an extra column of 1s (to account for the bias term $b_0$) - $w$ is a weight vector of shape
(14,)(13 feature weights + 1 bias term)
2. Rewrite the Gradient Descent Function
Here's the updated function that handles multi-feature data, plus we'll add feature scaling (critical for gradient descent to converge well with the Boston Dataset's varied feature scales):
import numpy as np from sklearn.datasets import load_boston import matplotlib.pyplot as plt from sklearn.model_selection import train_test_split from sklearn.linear_model import LinearRegression from sklearn.metrics import mean_squared_error from sklearn.preprocessing import StandardScaler # Load dataset dataset = load_boston() X = dataset.data y = dataset.target # Scale features (essential for stable gradient descent convergence) scaler = StandardScaler() X_scaled = scaler.fit_transform(X) # Add a column of 1s to X to include the bias term in our weight vector X_augmented = np.hstack([np.ones((X_scaled.shape[0], 1)), X_scaled]) def gradient_descent(X, y, eta=0.01, n_iter=1000): '''Gradient descent implementation for multi-feature linear regression eta: Learning Rate n_iter: Number of iterations ''' m = len(y) # Initialize weights: [b0, w1, w2, ..., w13] weights = np.zeros(X.shape[1]) cost_list = [] for i in range(n_iter): # Predict y using matrix multiplication (matches y's shape now) y_pred = X @ weights # Calculate error (valid since y and y_pred are both (506,)) error = y - y_pred # Compute gradient using matrix operations (efficient for any feature count) gradient = (2/m) * X.T @ error # Update weights weights -= eta * gradient # Store current MSE cost = mean_squared_error(y, y_pred) cost_list.append(cost) return cost_list, weights # Run gradient descent cost_history, final_weights = gradient_descent(X_augmented, y) # Plot cost reduction to verify convergence plt.plot(range(len(cost_history)), cost_history) plt.xlabel('Iterations') plt.ylabel('Mean Squared Error') plt.title('Cost Reduction During Gradient Descent') plt.show() # Compare with sklearn's LinearRegression for validation lr = LinearRegression() lr.fit(X_scaled, y) print("Our final bias (b0):", final_weights[0]) print("Sklearn's intercept:", lr.intercept_) print("\nOur feature weights:", final_weights[1:]) print("Sklearn's coefficients:", lr.coef_)
Key Changes Explained
- Feature Scaling: We used
StandardScalerto normalize features to mean 0 and variance 1. Without this, features with larger scales (likeRMorLSTAT) would dominate gradient updates, making convergence slow or unstable. - Augmented X: Adding a column of 1s lets us include the bias term in the weight vector, simplifying matrix operations.
- Weight Vector: Initialized
weightsas a zero vector matching the number of columns in the augmented X (14 elements: bias + 13 feature weights). - Matrix Gradient Calculation: Using
X.T @ errorcomputes the gradient efficiently for all weights at once, which works seamlessly for any number of features.
Why This Works
Now y_pred is a 1D array of shape (506,), matching y exactly—so the error calculation is valid. The gradient update uses linear algebra to compute the derivative of the MSE cost function for all weights simultaneously.
Bonus: Verify Convergence
The plotted cost history will show you if the MSE decreases steadily and plateaus, which means the algorithm is converging. If it oscillates or increases, tweak the learning rate (eta) or increase the number of iterations.
内容的提问来源于stack exchange,提问作者srbemr

