You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何tune::last_fit指标与summary()结果不同?哪个是测试集R²?

tune::collect_metrics()与summary()的R²差异解析

问题描述

使用独立数据集评估通过tune::last_fit()构建的模型时,发现tune::collect_metrics()得到的R²指标与summary()输出的R²结果不一致,需明确两者的计算区别,以及哪一个对应独立数据集观测值与预测值的R²。

可复现示例

library(recipes)
library(rsample)
library(parsnip)

set.seed(6735)
tr_te_split <- initial_split(mtcars)

spline_rec <- recipe(mpg ~ ., data = mtcars) %>%
  step_ns(disp)

lin_mod <- linear_reg() %>%
  set_engine("lm")

spline_res <- tune::last_fit(lin_mod, spline_rec, split = tr_te_split)

# 查看测试集性能指标
tune::collect_metrics(spline_res)
#> # A tibble: 2 × 4
#>   .metric .estimator .estimate .config             
#>   <chr>   <chr>          <dbl> <chr>                
#> 1 rmse    standard       3.80  Preprocessor1_Model1
#> 2 rsq     standard       0.729 Preprocessor1_Model1

# 查看训练集拟合模型的summary
spline_res %>% 
  parsnip::extract_fit_engine() %>% 
  summary()
#> 
#> Call:
#> stats::lm(formula = ..y ~ ., data = data)
#> 
#> 残差:
#>     Min      1Q  Median      3Q     Max 
#> -3.4453 -1.1980 -0.1464  1.3246  2.8223 
#> 
#> 系数:
#>               Estimate Std. Error t value Pr(>|t|)
#> (Intercept)  23.087028  18.641785   1.238    0.239
#> cyl           0.326218   1.402236   0.233    0.820
#> hp            0.005969   0.024848   0.240    0.814
#> drat         -0.009576   1.597293  -0.006    0.995
#> wt           -0.902839   2.503336  -0.361    0.725
#> qsec          0.185826   0.745021   0.249    0.807
#> vs            1.492756   2.255781   0.662    0.521
#> am            4.101555   3.110797   1.318    0.212
#> gear          0.174875   1.730223   0.101    0.921
#> carb         -1.278962   1.009824  -1.267    0.229
#> disp_ns_1   -15.149506  13.649995  -1.110    0.289
#> disp_ns_2    -4.905087   6.756046  -0.726    0.482
#> 
#> 残差标准误:2.397,自由度为12
#> 多重R平方: 0.9204,调整后R平方: 0.8473 
#> F统计量:12.61,自由度11和12,p值:5.869e-05

创建于2023-05-22,使用reprex v2.0.2

差异解析

1. summary()的R²:训练集拟合优度

extract_fit_engine()提取的是基于训练集拟合完成的lm模型,summary()输出的「多重R平方:0.9204」是模型在训练数据上的拟合优度,反映模型对训练样本的解释能力,属于内样本性能指标。

2. tune::collect_metrics()的R²:测试集预测性能

tune::last_fit()的执行逻辑是:用训练集拟合模型,然后自动在拆分好的测试集(独立数据集)上评估模型性能。collect_metrics()返回的R²(0.729)是模型在独立测试集上的预测R²,直接对应测试集观测值与预测值的相关平方,属于外样本泛化性能指标。

3. 手动验证测试集R²

可以通过提取测试集的预测值与真实值,手动计算R²来验证:

# 提取测试集的真实值和预测值
test_preds <- spline_res %>% extract_predictions()

# 手动计算测试集R²
cor(test_preds$mpg, test_preds$.pred)^2
# 输出结果≈0.729,与collect_metrics的结果一致

结论

  • 若关注模型在**独立数据集(测试集)**上的预测能力,应使用tune::collect_metrics()返回的R²。
  • summary()的R²仅反映模型在训练集上的拟合效果,不能代表模型的泛化能力。

内容的提问来源于stack exchange,提问作者Paul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 08:22:06