You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Sklearn LinearRegression预测结果及系数随特征列顺序变化的原因排查

Why does sklearn's LinearRegression give different predictions when feature column order changes?

Hey there! Let's dig into why you're seeing this weird behavior with LinearRegression. The core issue here is perfect multicollinearity in your feature data—let me break this down step by step.

What's causing the problem?

Look closely at your X_train dataframe: the values in columns x2 and x3 are identical for every single row. That means these two features are perfectly linearly correlated (their correlation coefficient is 1.0).

OLS linear regression assumes that the feature matrix has full rank—i.e., no features are exact linear combinations of others. When you have perfect multicollinearity, the design matrix becomes singular, and there's no unique solution for the regression coefficients.

Sklearn's LinearRegression uses singular value decomposition (SVD) under the hood to compute coefficients when the matrix is singular. The problem is that this decomposition can produce different "valid" coefficient sets depending on the order of features, and numerical precision issues can amplify this into drastically different predictions (like the huge e+12 values you saw in the first test).

Let's verify this

You can check the correlation between your features to confirm:

print(X_train.corr())

You'll see that x2 and x3 have a correlation of 1.0. You can also check the rank of the feature matrix:

from numpy.linalg import matrix_rank
print(matrix_rank(X_train))  # Outputs 2, but you have 3 features—proof of rank deficiency

How to fix this?

You have two main options:

  1. Remove duplicate/redundant features: Since x2 and x3 are identical, you can drop one of them entirely. This will make your feature matrix full rank, and the regression results will be consistent regardless of feature order.
    Example modified code:

    # Drop x3 from training and test data
    X_train_clean = X_train.drop('x3', axis=1)
    X_test_clean = X_test.drop('x3', axis=1)
    
    # Train model with clean data
    model_clean = LinearRegression().fit(X_train_clean, y_train)
    yhat_test_clean = model_clean.predict(X_test_clean)
    
    # Now sort columns and retrain
    cols_sorted = sorted(X_train_clean.columns)
    X_train_sorted = X_train_clean[cols_sorted]
    X_test_sorted = X_test_clean[cols_sorted]
    model_sorted_clean = LinearRegression().fit(X_train_sorted, y_train)
    yhat_test_sorted_clean = model_sorted_clean.predict(X_test_sorted)
    
    print(f'yhat_test_clean: {yhat_test_clean}')
    print(f'yhat_test_sorted_clean: {yhat_test_sorted_clean}')
    

    Now both predictions will be identical.

  2. Use a regularized regression model: If you need to keep all features (though in this case they're identical, so it's unnecessary), models like RidgeRegression add a regularization penalty that stabilizes the coefficient estimates even with multicollinearity. This will also give consistent results regardless of feature order.

Key takeaway

OLS regression requires full-rank features to have a unique, stable solution. Perfect multicollinearity breaks this assumption, leading to unstable, order-dependent results. Fixing the feature redundancy will resolve the issue.

内容的提问来源于stack exchange,提问作者bugfoot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 23:12:47