在R中按Subject与Conditions分组计算Score列的均值
按Subject和Conditions分组计算Score均值的方法
给定如下R数据框:
example <- data.frame( Subject = c(rep(101, 8), rep(102, 8)), Run = c(1,1,1,1,2,2,2,2,1,1,1,1,2,2,2,2), Conditions = c(rep(c('SIN', 'SIN', 'SRN', 'SRN'), 4)), Score = c(1,1,1,1,2,2,2,2,3,3,3,3,4,4,4,4) )
需要按Subject和Conditions分组,计算Score列的均值,期望输出:
solution <- data.frame( Subject = c(101, 101, 102, 102), Conditions = c(rep(c('SIN', 'SRN'),2)), Score = c(1.5,1.5,3.5,3.5) )
以下是两种常用实现方法:
Base R 实现
使用aggregate()函数直接完成分组计算:
result_base <- aggregate(Score ~ Subject + Conditions, data = example, FUN = mean)
- 公式
Score ~ Subject + Conditions定义了分组依据(Subject和Conditions)与待计算的变量(Score) FUN = mean指定对分组后的Score执行均值计算
dplyr 实现
如果习惯使用tidyverse风格的语法,可通过dplyr包的group_by()和summarize()组合实现:
library(dplyr) result_dplyr <- example %>% group_by(Subject, Conditions) %>% summarize(Score = mean(Score), .groups = "drop")
group_by(Subject, Conditions)指定分组变量summarize(Score = mean(Score))计算每组的均值并保留列名.groups = "drop"用于取消分组状态,返回普通数据框结构
结果验证
执行上述代码后,可通过以下命令验证结果是否与目标一致:
all.equal(result_base, solution) # 返回 TRUE all.equal(result_dplyr, solution) # 返回 TRUE
内容的提问来源于stack exchange,提问作者jo_
相关产品推荐
相关产品推荐

