为R数据集创建循环/函数,按多维度分区域计算结果占比
R数据集分组占比计算方案
需求明确
需要针对数据集里的outcome1、outcome2、outcome3三个结果变量,分别按sex、age_group、income、education这四个维度计算各类别的占比;同时要按照geography字段的area1和area2生成两个独立的结果输出。
改进后的代码实现
library(tidyverse) df <- data.frame ( outcome1 = c("poor", "good", "excellent", "poor", "good", "poor", "poor", "excellent"), outcome2 = c("good", "excellent", "excellent", "poor", "excellent", "poor", "excellent", "poor"), outcome3 = c("poor", "poor", "excellent", "poor", "poor", "poor", "excellent", "good"), sex = c("F", "M", "M", "F", "F", "M", "F", "M"), age_group = c("50-54", "60-64", "80+", "70-74", "40-44", "45-49", "60-64", "65-69"), income = c("$<40,000", "$50,000-79,000", "$80,000-110,000", "$111,000+", "$<40,000", "$<40,000", "$50,000-79,000", "$80,000-110,000"), education = c("HS", "College", "Bachelors", "Masters", "HS", "College", "Bachelors", "Masters"), geography= c("area1", "area2", "area1", "area2", "area2", "area1", "area2", "area1") ) # 定义占比计算函数 calc_proportion <- function(outcome_col, group_col, data) { data %>% group_by(geography, {{outcome_col}}, {{group_col}}) %>% summarise(count = n(), .groups = "drop_last") %>% mutate(total = sum(count), proportion = count / total * 100) %>% ungroup() } # 指定需要处理的变量组合 outcomes <- c("outcome1", "outcome2", "outcome3") groups <- c("sex", "age_group", "income", "education") # 批量计算所有组合的占比 all_results <- crossing(outcome = outcomes, group = groups) %>% mutate(result = map2(outcome, group, ~calc_proportion(.x, .y, df))) %>% unnest(result) # 拆分出area1和area2的独立结果 area1_results <- all_results %>% filter(geography == "area1") area2_results <- all_results %>% filter(geography == "area2") # 查看结果示例 print(area1_results %>% head()) print(area2_results %>% head())
代码说明
- 函数复用:
calc_proportion函数可适配任意结果变量和分组维度,自动按地区、结果、分组维度统计数量,计算每组占对应地区-结果组合的总比例,避免了硬编码总数的错误。 - 批量处理:用
crossing生成所有结果变量与分组维度的组合,通过map2批量调用计算函数,一次性完成所有统计需求。 - 结果拆分:最后按
geography筛选出两个地区的独立结果,方便后续分析或导出。
内容的提问来源于stack exchange,提问作者R_coder_new
相关产品推荐
相关产品推荐

