如何在R中基于数据框排名准则执行多数投票筛选最优模型
基于多数投票筛选最优预测模型的R实现
场景说明
现有包含8个预测模型评估指标的数据框,其中:
- 误差指标(RMSE、MAE):数值越小,模型性能越好
- 相关性指标(R²、CCC):数值越大,模型性能越好
需要通过每个指标筛选出最优模型,再经多数投票确定最终最优模型。
步骤实现
1. 准备数据与依赖包
首先加载tidyverse工具包,并导入模型评估数据:
library(tidyverse) # 模型评估数据框 dat <- structure(list(model_name = c("Random Forest", "XGBoost", "XGBoost-reg", "Null model", "Plain LM", "Elastic LM", "LM-pep.charge", "LM-rf.10vip"), RMSE = c(0.853, 0.886, 0.719, 2.41, 16.6, 0.731, 1.16, 1.03), MAE = c(0.545, 0.708, 0.589, 1.98, 8.6, 0.588, 0.874, 0.729), `R^2` = c(0.806, 0.865, 0.915, NA, 0.0645, 0.927, 0.8, 0.822), ccc = c(0.89, 0.928, 0.951, 0, 0.0685, 0.945, 0.847, 0.901)), row.names = c(NA, -8L), class = c("tbl_df", "tbl", "data.frame"))
2. 定义指标优化方向
明确每个指标的最优判断规则:
# 定义指标优化方向:min=越小越好,max=越大越好 metric_direction <- c( RMSE = "min", MAE = "min", `R^2` = "max", ccc = "max" )
3. 筛选各指标的最优模型
将数据转为长格式,按指标分组筛选出每个指标对应的最优模型:
# 找出每个指标的最优模型 top_models <- dat %>% pivot_longer(cols = -model_name, names_to = "metric", values_to = "value") %>% filter(!is.na(value)) %>% # 剔除含NA的无效数据 group_by(metric) %>% mutate( # 根据指标方向标记最优模型 is_top = case_when( metric_direction[metric] == "min" ~ value == min(value, na.rm = TRUE), metric_direction[metric] == "max" ~ value == max(value, na.rm = TRUE) ) ) %>% filter(is_top) %>% ungroup() %>% select(metric, model_name) # 查看各指标的最优模型结果 print(top_models)
运行后会得到和手动筛选一致的结果:
# A tibble: 4 × 2 metric model_name <chr> <chr> 1 RMSE XGBoost-reg 2 MAE Random Forest 3 R^2 Elastic LM 4 ccc XGBoost-reg
4. 多数投票确定最终最优模型
统计每个模型的得票数,选出得票最高的模型:
# 统计得票数并选出最优 winner <- top_models %>% count(model_name, sort = TRUE) %>% slice(1) # 输出结果 cat("多数投票选出的最优模型:", winner$model_name, "\n得票数:", winner$n, "\n")
最终输出:
多数投票选出的最优模型: XGBoost-reg 得票数: 2
内容的提问来源于stack exchange,提问作者littleworth
相关产品推荐
相关产品推荐

