如何利用StatsModels OLS获取测试数据的回归结果摘要(含AIC、R-squared)
Got it, let's break this down for you. You've already built your OLS model with StatsModels and can make predictions on new data, but now you want metrics like R-squared, AIC, etc., specifically for your test set. There are two common approaches here—let's cover both, since they serve different purposes:
Option 1: Evaluate Your Trained Model on the Test Set (Most Common)
This is what you probably want: using your already-trained model to calculate how well it performs on the test data, including R-squared, AIC, MSE, etc. You don't need to re-fit the model—just use the predictions from your existing model to compute these metrics.
Here's a step-by-step code example:
import statsmodels.api as sm import statsmodels.tools.eval_measures as em from sklearn.metrics import r2_score # Or use statsmodels' built-in r2_score # Assume you already have: # - Your trained OLS model: `model = sm.OLS(y_train, X_train).fit()` # - Test data: `X_test` (features) and `y_test` (actual NOx concentrations) # Important: If you added a constant term to your training data, add it to X_test too! X_test = sm.add_constant(X_test) # Get predictions for the test set y_pred = model.predict(X_test) # Calculate R-squared (measures how well predictions fit actual values) test_r2 = r2_score(y_test, y_pred) # Alternatively, use StatsModels' version: test_r2 = em.r2_score(y_test, y_pred) # Calculate AIC (Akaike Information Criterion) # AIC formula: -2 * log_likelihood + 2 * number_of_parameters test_log_likelihood = model.loglike(y_test, y_pred) num_params = model.params.shape[0] # Includes intercept if you added it test_aic = -2 * test_log_likelihood + 2 * num_params # Bonus: Calculate other useful metrics like MSE (Mean Squared Error) test_mse = em.mse(y_test, y_pred) # Print out the results print(f"Test Set R-squared: {test_r2:.4f}") print(f"Test Set AIC: {test_aic:.4f}") print(f"Test Set MSE: {test_mse:.4f}")
Option 2: Re-Fit the Model on the Test Set (Rarely Needed)
If for some reason you want a full OLS summary (like the one you get from model.summary() for training data) but based on fitting the model only on the test set, you can re-fit the model with your test data. Note: This isn't standard model evaluation—this creates an entirely new model trained on test data, which is usually not what you want.
But if you need it, here's how:
# Re-fit OLS on test data test_model = sm.OLS(y_test, X_test).fit() # Print the full summary (includes AIC, R-squared, coefficients, p-values, etc.) print(test_model.summary())
Quick Note
Double-check that your X_test has the same structure as your training X_train—especially if you added a constant term during training (using sm.add_constant). If you forget this, your predictions and metrics will be incorrect.
内容的提问来源于stack exchange,提问作者eliavs

