You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中不使用rstatix包对分组变量批量执行T检验

按物品分组执行独立样本T检验(无rstatix依赖)

需求说明

现有包含物品、成本和分组的数据框,需要针对每个物品,检验两组成本均值是否存在差异,要求不使用rstatix包,优先用base R的lapply/循环,也可使用tidyr和dplyr。

示例数据

df = structure(list(Item = structure(c(1L, 1L, 1L, 1L, 2L, 2L, 2L, 
2L, 3L, 3L, 3L, 3L, 3L, 4L, 4L, 4L, 4L), .Label = c("Book A", 
"Book B", "Book C", "Book D"), class = "factor"), Cost = c(7L, 
9L, 6L, 7L, 4L, 6L, 5L, 3L, 5L, 4L, 7L, 2L, 2L, 4L, 2L, 9L, 4L
), Grouping = structure(c(1L, 2L, 1L, 2L, 1L, 2L, 1L, 1L, 2L, 
1L, 1L, 2L, 2L, 1L, 2L, 2L, 1L), .Label = c("A", "B"), class = "factor")), class = "data.frame", row.names = c(NA, 
-17L))

方法1:Base R 实现(lapply)

通过split()按物品拆分数据框,再用lapply()遍历每个子集执行t检验,最后提取p值并整理成结果表格:

# 按Item拆分数据框
split_df <- split(df, df$Item)

# 遍历每个子集,执行t检验并提取p值
p_values <- sapply(split_df, function(x) {
  # 确保两组都有数据(避免单组报错)
  if(length(unique(x$Grouping)) == 2) {
    t.test(Cost ~ Grouping, data = x)$p.value
  } else {
    NA  # 若某物品只有一组数据,返回NA
  }
})

# 整理成目标格式的数据框
result_base <- data.frame(
  Item = names(p_values),
  `P-Value (H0: Mean of group A = Mean of group B)` = round(p_values, 4),
  stringsAsFactors = FALSE
)

print(result_base)

方法2:dplyr + tidyr 实现

利用tidyverse的分组和汇总功能,在分组内执行t检验并提取p值:

library(dplyr)

result_tidy <- df %>%
  group_by(Item) %>%
  summarize(
    `P-Value (H0: Mean of group A = Mean of group B)` = {
      if(n_distinct(Grouping) == 2) {
        t.test(Cost ~ Grouping, data = cur_data())$p.value
      } else {
        NA
      }
    },
    .groups = "drop"
  ) %>%
  mutate(across(where(is.numeric), ~round(., 4)))

print(result_tidy)

输出结果

两种方法都会得到如下格式的结果:

ItemP-Value (H0: Mean of group A = Mean of group B)
Book A0.6927
Book B0.4677
Book C0.1217
Book D0.3033

内容的提问来源于stack exchange,提问作者Luther_Proton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 09:54:23