R语言:按diet和period分组求多列均值的替代方法咨询
按diet和period分组计算多列均值的多种R实现方法
嘿,我来给你整理几种除了aggregate()之外的实现思路,都是R里常用的分组计算均值的手段,先把你提供的测试数据代码放这里,方便大家复现:
set.seed(8) id <- 1:6 diet <- rep(c("A","B"),3) period <- rep(c(1,2),3) score1 <- sample(1:100,6) score2 <- sample(1:100,6) score3 <- sample(1:100,6) df <- data.frame(id,diet,period,score1, score2, score3)
方法1:使用dplyr(tidyverse生态)
dplyr是tidyverse系列里的核心包,语法非常直观,适合做数据清洗和分组统计。我们可以用group_by()指定分组变量,再用summarise(across(...))批量处理多个score列:
library(dplyr) df_result <- df %>% group_by(period, diet) %>% summarise(across(starts_with("score"), mean), .groups = "drop") # 如果需要把period作为行名(和你期望的输出格式更匹配) df_result <- df_result %>% arrange(period) %>% column_to_rownames("period") %>% select(-diet) # 可根据需求保留/移除diet列
方法2:使用data.table包
data.table处理大规模数据集时效率极高,语法简洁紧凑,适合快速做分组运算:
library(data.table) dt <- as.data.table(df) dt_result <- dt[, .(score1 = mean(score1), score2 = mean(score2), score3 = mean(score3)), by = .(period, diet)] # 调整格式匹配需求 dt_result <- dt_result[order(period), ] rownames(dt_result) <- dt_result$period dt_result <- dt_result[, !c("period", "diet")]
方法3:Base R的tapply + do.call组合
不用任何第三方包,纯Base R就能实现。我们可以用tapply对每个score列按分组计算均值,再用do.call把结果合并成数据框:
# 创建分组变量 group_var <- interaction(df$period, df$diet) # 对每个score列计算分组均值 score1_mean <- tapply(df$score1, group_var, mean) score2_mean <- tapply(df$score2, group_var, mean) score3_mean <- tapply(df$score3, group_var, mean) # 合并成数据框并整理格式 base_result <- do.call(cbind, list(score1 = score1_mean, score2 = score2_mean, score3 = score3_mean)) base_result <- as.data.frame(base_result) # 提取period作为行名并排序 base_result$period <- as.integer(sapply(strsplit(rownames(base_result), "\\."), `[`, 1)) base_result <- base_result[order(base_result$period), ] rownames(base_result) <- base_result$period base_result <- base_result[, !"period"]
方法4:Base R的by函数
by函数可以按分组对数据框进行操作,返回的结果可以再整合成数据框:
# 按period和diet分组,对每个分组计算score列的均值 by_result <- by(df[, c("score1", "score2", "score3")], list(period = df$period, diet = df$diet), colMeans) # 把by返回的列表转换成数据框 by_result <- do.call(rbind, by_result) by_result <- as.data.frame(by_result) # 调整行名和排序 by_result$period <- as.integer(sapply(strsplit(rownames(by_result), "\\."), `[`, 1)) by_result <- by_result[order(by_result$period), ] rownames(by_result) <- by_result$period by_result <- by_result[, !"period"]
运行以上任意一种方法,最终都能得到你期望的输出格式:
score1 score2 score3 1 52.33333 50 19.66667 2 51.33333 55 56.66667
内容的提问来源于stack exchange,提问作者user11916948
相关产品推荐
相关产品推荐

