嵌套列映射遍历及两组t检验计算的tidy R实现方法咨询
实现方案
基于tidyverse生态的purrr+dplyr+broom组合即可完成批量计算,全程符合tidy风格编码规范:
1. 依赖包加载
library(tidyverse) library(broom)
2. 基础批量实现
直接遍历所有嵌套tibble完成t检验,支持保留完整检验对象或直接输出结构化统计结果:
# 方案1:仅计算并保留完整t检验对象 result_with_raw_test <- data %>% mutate( t_test = map(data, function(df) { grp1 <- df %>% filter(x == 1) %>% pull(y) grp2 <- df %>% filter(x == 5) %>% pull(y) t.test(grp1, grp2) }) ) # 方案2:直接输出结构化统计结果(推荐,方便后续分析) result_tidy <- data %>% mutate( # 用公式写法更简洁,无需单独提取两组向量 t_test = map(data, ~t.test(y ~ x, data = filter(.x, x %in% c(1,5)))), # 将t检验返回的列表转为标准tibble格式 test_stat = map(t_test, tidy) ) %>% # 展开检验结果,每一行对应一个日期的t检验统计指标 unnest(test_stat) %>% # 可选:按需保留需要的列 select(date, estimate, statistic, p.value, conf.low, conf.high, method)
3. 增强健壮性的实现
如果存在某组样本量不足的情况,可增加校验逻辑避免报错:
result_safe <- data %>% mutate( test_stat = map(data, function(df) { df_sub <- filter(df, x %in% c(1,5)) # 校验两组都至少有2个有效样本才执行t检验 if(n_distinct(df_sub$x) == 2 && min(table(df_sub$x)) >= 2) { t.test(y ~ x, data = df_sub) %>% tidy() } else { # 样本不足时返回空tibble,后续自动填充NA tibble() } }) ) %>% unnest(test_stat, keep_empty = TRUE)
内容的提问来源于stack exchange,提问作者user113156
相关产品推荐
相关产品推荐

