如何在R语言中使用nest/nest_by后添加均值列?报错求解
解决方案
错误分析
你遇到的问题主要有三个:
- 自定义的
my_mean函数没有引用传入的data参数,直接调用score会导致变量未找到的错误,因为score是嵌套数据集里的列,必须通过data对象访问。 - 多余的
unite操作将condition和time合并,生成了pre_0、post_1、post_2三个组合,而非你需要的按time(0、1、2)的三个时间段分组。 - 整体均值的计算时机错误:你在嵌套前的
mutate(score2 = mean(score))会因为后续的nest_by分组,导致score2变成组内均值而非全局整体均值。
方案1:使用nest()
library(dplyr) library(purrr) library(tidyr) set.seed(1414) test <- tibble(id = c(1:100), condition = c(rep(c("pre", "post"), 50)), time = c(case_when(condition == "pre" ~ 0, condition == "post" ~ sample(c(1, 2), size = 100, replace = TRUE))), score = case_when(time == 0 ~ 1, time == 1 ~ 10, time == 2 ~ 100)) # 1. 先计算全局整体均值,添加到每一行 test_with_global_mean <- test %>% mutate(global_mean = mean(score, na.rm = TRUE)) # 2. 按time分组嵌套,得到每个时间段的独立数据集 nested_data <- test_with_global_mean %>% nest(data = -c(time, global_mean)) # 3. 计算每个嵌套数据集的组均值 nested_data <- nested_data %>% mutate(group_mean = map_dbl(data, ~ mean(.x$score, na.rm = TRUE))) # 查看结果 nested_data
方案2:使用nest_by()
nest_by()返回行式 tibble,无需额外用map,可以直接在mutate中操作每组的data:
# 先计算全局整体均值 global_mean_val <- mean(test$score, na.rm = TRUE) nested_by_data <- test %>% nest_by(time) %>% # 添加全局整体均值列 mutate(global_mean = global_mean_val, # 计算组均值 group_mean = mean(data$score, na.rm = TRUE)) # 查看结果 nested_by_data
关键说明
- 全局整体均值需要提前计算(或在嵌套前添加到所有行),确保每个分组的
global_mean都是同一个全局值。 - 计算组均值时,必须明确引用嵌套数据集中的
score列(如data$score或.x$score),避免变量找不到的错误。 - 两种方案都保留了每个
time对应的独立数据集(data列),满足你后续计算的需求。
内容的提问来源于stack exchange,提问作者J.Sabree
相关产品推荐
相关产品推荐

