Sklearn LinearRegression预测结果及系数随特征列顺序变化的原因排查
Hey there! Let's dig into why you're seeing this weird behavior with LinearRegression. The core issue here is perfect multicollinearity in your feature data—let me break this down step by step.
What's causing the problem?
Look closely at your X_train dataframe: the values in columns x2 and x3 are identical for every single row. That means these two features are perfectly linearly correlated (their correlation coefficient is 1.0).
OLS linear regression assumes that the feature matrix has full rank—i.e., no features are exact linear combinations of others. When you have perfect multicollinearity, the design matrix becomes singular, and there's no unique solution for the regression coefficients.
Sklearn's LinearRegression uses singular value decomposition (SVD) under the hood to compute coefficients when the matrix is singular. The problem is that this decomposition can produce different "valid" coefficient sets depending on the order of features, and numerical precision issues can amplify this into drastically different predictions (like the huge e+12 values you saw in the first test).
Let's verify this
You can check the correlation between your features to confirm:
print(X_train.corr())
You'll see that x2 and x3 have a correlation of 1.0. You can also check the rank of the feature matrix:
from numpy.linalg import matrix_rank print(matrix_rank(X_train)) # Outputs 2, but you have 3 features—proof of rank deficiency
How to fix this?
You have two main options:
Remove duplicate/redundant features: Since
x2andx3are identical, you can drop one of them entirely. This will make your feature matrix full rank, and the regression results will be consistent regardless of feature order.
Example modified code:# Drop x3 from training and test data X_train_clean = X_train.drop('x3', axis=1) X_test_clean = X_test.drop('x3', axis=1) # Train model with clean data model_clean = LinearRegression().fit(X_train_clean, y_train) yhat_test_clean = model_clean.predict(X_test_clean) # Now sort columns and retrain cols_sorted = sorted(X_train_clean.columns) X_train_sorted = X_train_clean[cols_sorted] X_test_sorted = X_test_clean[cols_sorted] model_sorted_clean = LinearRegression().fit(X_train_sorted, y_train) yhat_test_sorted_clean = model_sorted_clean.predict(X_test_sorted) print(f'yhat_test_clean: {yhat_test_clean}') print(f'yhat_test_sorted_clean: {yhat_test_sorted_clean}')Now both predictions will be identical.
Use a regularized regression model: If you need to keep all features (though in this case they're identical, so it's unnecessary), models like
RidgeRegressionadd a regularization penalty that stabilizes the coefficient estimates even with multicollinearity. This will also give consistent results regardless of feature order.
Key takeaway
OLS regression requires full-rank features to have a unique, stable solution. Perfect multicollinearity breaks this assumption, leading to unstable, order-dependent results. Fixing the feature redundancy will resolve the issue.
内容的提问来源于stack exchange,提问作者bugfoot

