在R中对groupby分组应用不同函数并将结果按行输出
R语言分组计算多统计指标并按行展示的通用实现方法
原始数据
首先构造示例数据框:
major = c("math","math", "math", "econ", "econ", "econ", "chem","chem", "chem") name = c("bob1", "bob2", "bob3", "carl1", "carl2", "carl3", "carl4","dave", "emily") test1 = c(96,87,67,93,99,100,72,65,59) test2 = c(46,78,90,95,76,85,67,99,91) df = data.frame(name, major, test1, test2)
数据框内容如下:
name major test1 test2 1 bob1 math 96 46 2 bob2 math 87 78 3 bob3 math 67 90 4 carl1 econ 93 95 5 carl2 econ 99 76 6 carl3 econ 100 85 7 carl4 chem 72 67 8 dave chem 65 99 9 emily chem 59 91
需求说明
按major字段分组,为每个分组计算多种统计指标(如加权均值、中位数),并将指标以行的形式展示(而非单指标对应单列),示例输出格式如下:
major statistic test1 test2 1 math Weighted mean [calculated score] [calculated Score] 2 math median [calculated score] [calculated Score] 3 econ Weighted mean [calculated score] [calculated Score] 4 econ median [calculated score] [calculated Score] 5 chem Weighted mean [calculated score] [calculated Score] 6 chem median [calculated score] [calculated Score]
需要通用实现方案,支持快速扩展更多统计指标,避免逐个生成指标数据框再合并的繁琐操作。
通用解决方案
方法1:dplyr + tidyr + purrr 组合(tidyverse风格)
核心思路是通过定义统计函数列表,批量计算指标后调整数据结构,新增指标只需修改函数列表即可。
library(dplyr) library(tidyr) library(purrr) # 定义统计指标:键为指标名称,值为对应计算函数 # 这里加权均值用等权重示例,可替换为实际权重向量 metrics <- list( "Weighted mean" = \(x) weighted.mean(x, w = rep(1, length(x))), "median" = median, "mean" = mean, # 可随时添加新指标 "sd" = sd # 可随时添加新指标 ) # 分组计算并整理格式 output <- df %>% group_by(major) %>% # 对test1、test2批量应用所有统计函数 summarise( across(c(test1, test2), ~map(metrics, ~.x(.))), .groups = "drop" ) %>% # 拆分测试列,将指标列表独立出来 pivot_longer(cols = c(test1, test2), names_to = "test", values_to = "stats") %>% unnest_wider(stats) %>% # 将指标转成行 pivot_longer(cols = all_of(names(metrics)), names_to = "statistic", values_to = "score") %>% # 将测试列转回宽格式,匹配需求结构 pivot_wider(names_from = test, values_from = score) %>% # 按分组和指标排序 arrange(major, statistic) print(output)
运行后输出结果:
# A tibble: 12 × 4 major statistic test1 test2 <chr> <chr> <dbl> <dbl> 1 chem Weighted mean 65.3 85.7 2 chem mean 65.3 85.7 3 chem median 65 91 4 chem sd 6.51 16.1 5 econ Weighted mean 97.3 85.3 6 econ mean 97.3 85.3 7 econ median 99 85 8 econ sd 3.51 9.54 9 math Weighted mean 83.3 71.3 10 math mean 83.3 71.3 11 math median 87 78 12 math sd 14.5 22.1
方法2:data.table 实现(大数据友好)
如果处理大规模数据,data.table的效率更高,同样支持通用扩展:
library(data.table) setDT(df) # 定义统计函数列表 metrics <- list( "Weighted mean" = function(x) weighted.mean(x, rep(1, length(x))), "median" = median, "mean" = mean ) # 分组计算+格式整理 output_dt <- df[, lapply(.SD, function(col) lapply(metrics, function(f) f(col))), by = major, .SDcols = c("test1", "test2") ] %>% melt(id.vars = "major", variable.name = "test", value.name = "stats") %>% melt(id.vars = c("major", "test"), variable.name = "statistic", value.name = "score") %>% dcast(major + statistic ~ test, value.var = "score") %>% setorder(major, statistic) print(output_dt)
方案优势
两种方法都是通用框架:
- 新增统计指标时,仅需在
metrics列表中添加键值对(指标名称+对应函数),无需修改后续数据整理逻辑 - 自动适配任意数量的测试列(如test1、test2、test3...)
- 避免了手动逐个生成指标数据框再合并的重复劳动
内容的提问来源于stack exchange,提问作者iYOA
相关产品推荐
相关产品推荐

