R语言中基于两个变量递进筛选dataframe自动生成统计表格的方法
实现方案
核心思路是先构造两个变量所有递进阈值的交叉组合,批量计算每个组合对应的统计值后,直接转换为你需要的行列结构的表格,无需逐个手动写filter语句。
完整可运行代码
# 加载依赖包 library(tidyverse) # 构造示例数据集 df <- tibble( age = c(45, 50, 55, 60, 65), risk = c(5, 45, 70, 80, 50), disease = c(1,0,1,1,0) ) # 定义递进筛选阈值,也可通过seq函数自动生成 # age阈值示例:也可以写为 seq(40, max(df$age), 5) 自动按5岁步长生成所有阈值 age_thresholds <- c(40, 45, 50, 55, 60) # risk阈值示例:也可以写为 seq(20, max(df$risk), 5) 自动按5分步长生成所有阈值 risk_thresholds <- c(20, 25, 30, 40, 50, 60, 70, 80) # 批量计算后转成目标表格 result <- crossing(age_thres = age_thresholds, risk_thres = risk_thresholds) %>% rowwise() %>% summarise( age_filter = paste0("age>", age_thres), risk_filter = paste0("risk>", risk_thres), # 计算患病率,需要统计满足条件的计数时,替换为 count = sum(df$age > age_thres & df$risk > risk_thres) prop_dis = sum(df$disease[df$age > age_thres & df$risk > risk_thres]) / sum(df$age > age_thres & df$risk > risk_thres), .groups = "drop" ) %>% # 转换为宽表:行是risk筛选条件,列是age筛选条件 pivot_wider( id_cols = risk_filter, names_from = age_filter, values_from = prop_dis ) # 查看生成的表格 print(result)
补充说明
- 结果中出现的
NaN代表对应行列的筛选组合下没有符合条件的观测,可以通过replace_na函数批量替换为0或者其他自定义值 - 阈值支持动态生成,不需要手动逐个输入,比如age从40开始每3岁递增到最大年龄可以写为
age_thresholds <- seq(40, max(df$age), 3)
内容的提问来源于stack exchange,提问作者Pras_V
相关产品推荐
相关产品推荐

