You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

因变量波动极小时回归模型性能评估:R²指标适用性探讨

Answers to Your Regression Metric Questions

Question 1: What metrics should use when the dependent variable has minimal variation?

When your target variable has extremely low variance (like your example where changes only show up in the second decimal place), R² becomes highly misleading—it measures the proportion of variance explained, and when there’s barely any variance to begin with, even a trivial model can score near 1.0.

Instead, focus on metrics that measure absolute or relative prediction error directly, which better reflect real-world performance:

  • Mean Absolute Error (MAE): Tells you the average absolute difference between predictions and true values. It’s intuitive and robust to outliers, making it easy to grasp how far off your predictions are on average.
  • Mean Absolute Percentage Error (MAPE): Useful for understanding error relative to the magnitude of the target. Since your values cluster around 102, MAPE will show you how big your errors are in percentage terms, helping you judge if they’re acceptable for your use case.
  • Mean Squared Error (MSE)/Root Mean Squared Error (RMSE): Penalizes larger errors more heavily than MAE. Use this if avoiding big mistakes is a priority for your model.
  • Mean Absolute Scaled Error (MASE): Compares your model’s error to a simple baseline (like predicting the overall target mean or the previous value). This helps you confirm your model is actually adding value over a naive approach.

Question 2: Is R² appropriate here, and does a 99% score mean great performance?

First, a critical note on your code: you’ve got the arguments for r2_score reversed! Scikit-learn expects true values first, then predictions:

score_test = r2_score(y_test, y_pred)  # Correct order

If you used the reversed order, your 99% score might not be accurate—so double-check that first.

Assuming the score is valid, R² is not a good choice here, and that high score doesn’t guarantee great performance. Here’s why:
R² is calculated as:

R² = 1 - (Sum of Squared Residuals / Total Sum of Squares)

Your target variable has a tiny total sum of squares (SS_tot) because the values barely vary. Even if your model makes consistent small errors, the ratio of residual sum of squares (SS_res) to SS_tot will be tiny, pushing R² close to 1.0. But those "small" errors could be meaningful—for example, if this is a precision sensor reading, a 0.01 error might be critical for your application.

To get a true sense of your model’s performance:

  1. Switch to one of the error metrics mentioned above (MAE, MAPE, etc.) to see the actual size of your predictions’ mistakes.
  2. Compare your model’s performance to a naive baseline (like predicting the mean of the target for all samples). If your model only slightly outperforms this baseline, it’s not as strong as the R² score suggests.

内容的提问来源于stack exchange,提问作者chink

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:46:48