使用PyTorch构建多元线性回归模型,结果不符合预期求解决
Let's break down why your linear regression model isn't performing as expected and fix it step by step.
Key Issues in Your Code
Missing Feature Normalization/Standardization
The Boston Housing dataset features have wildly different scales (e.g.,CRIMranges from ~0 to 88, whileRMsits between ~3 and 9). Without scaling, SGD struggles to converge—large-scale features will dominate gradient updates, making it nearly impossible for the model to learn meaningful weights for smaller-scale features.Inconsistent Tensor Shapes
Yourtargetsare 1D tensors ([N]), but the model'soutputsare 2D ([N, 1]). While PyTorch's MSELoss might handle this via broadcasting, misaligned shapes can lead to unexpected loss calculations or silent errors.Redundant Data Loading
You're converting numpy arrays to tensors inside the training loop—this is inefficient and unnecessary. Do this preprocessing once before starting training.Insufficient Training Epochs & Poor Learning Rate Choice
10 epochs is way too few for SGD to converge, especially with unnormalized data. A fixed learning rate of 0.01 is also risky: it might be too large (causing loss to explode) or too small (leading to no meaningful progress) when features aren't scaled.No Train/Test Split
Training on the entire dataset means you can't evaluate how well the model generalizes to unseen data—you have no way to tell if it's learning patterns or just memorizing the training set.
Fixed Implementation
Here's the corrected code addressing all the above issues:
from sklearn.datasets import fetch_california_housing # Boston dataset is deprecated, use this instead from sklearn.preprocessing import StandardScaler from sklearn.model_selection import train_test_split import torch import torch.nn as nn # Load dataset (ethical concerns led to Boston dataset being removed from scikit-learn) housing = fetch_california_housing() X, y = housing.data, housing.target # Split data into training and test sets X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) # Standardize features to mean=0, std=1 (critical for gradient-based optimizers) scaler = StandardScaler() X_train_scaled = scaler.fit_transform(X_train) X_test_scaled = scaler.transform(X_test) # Convert to PyTorch tensors and align target shape with model output X_train_tensor = torch.tensor(X_train_scaled, dtype=torch.float32) y_train_tensor = torch.tensor(y_train, dtype=torch.float32).unsqueeze(1) # Shape becomes [N, 1] X_test_tensor = torch.tensor(X_test_scaled, dtype=torch.float32) y_test_tensor = torch.tensor(y_test, dtype=torch.float32).unsqueeze(1) # Initialize model, loss function, and optimizer input_size = X_train.shape[1] linear_model = nn.Linear(input_size, 1, bias=True) criterion = nn.MSELoss() optimizer = torch.optim.SGD(linear_model.parameters(), lr=0.01) # LR works better with scaled data # Alternatively, use Adam for more stable convergence: optimizer = torch.optim.Adam(linear_model.parameters(), lr=0.001) num_epochs = 1000 # More epochs to allow convergence # Training loop for epoch in range(num_epochs): # Forward pass outputs = linear_model(X_train_tensor) loss = criterion(outputs, y_train_tensor) # Backward pass and optimize optimizer.zero_grad() loss.backward() optimizer.step() # Print progress at intervals if (epoch + 1) % 100 == 0: print(f'Epoch [{epoch+1}/{num_epochs}], Loss: {loss.item():.4f}') # Evaluate model on unseen test data with torch.no_grad(): test_outputs = linear_model(X_test_tensor) test_loss = criterion(test_outputs, y_test_tensor) print(f'\nTest Loss: {test_loss.item():.4f}')
Additional Notes
- Dataset Replacement: The Boston Housing dataset was removed from scikit-learn due to ethical issues, so we use the California Housing dataset as a suitable alternative.
- Normalization:
StandardScalerensures all features contribute equally to gradient updates, which is essential for SGD to converge properly. - Shape Alignment:
unsqueeze(1)converts the 1D target tensor to 2D, matching the model's output shape and eliminating broadcasting ambiguity. - Training Stability: Increasing epochs and using optimizers like Adam (more robust than vanilla SGD) helps the model reach a stable loss minimum.
- Evaluation: The test set lets you validate that your model is learning generalizable patterns, not just overfitting to training data.
内容的提问来源于stack exchange,提问作者saluo

