如何在R中针对配对病例-对照数据对多变量执行配对t检验
嘿,看起来你在处理配对病例-对照数据的多变量t检验时遇到了卡点,我来帮你理清楚正确的操作步骤,应该能解决你的问题~
先把你的问题场景和数据情况贴出来:
My figures are roughly as follows:
outcome paired ID variable A variable B variable C ... ...variable 99
1 1 0.163928565 1.384097885 2.590867531 1.810971177
0 1 0.351326004 1.231992145 3.309521365 1.838901895
1 2 0.290323902 1.224666519 3.346976812 1.804317302
0 2 0.086592296 0.967414001 3.127342374 1.95104196
0 3 0.222788134 1.650655928 2.980752068 2.186048898
1 3 0.313679639 0.740220592 3.171875959 2.006938751
1 4 0.053655263 1.520590707 3.122244134 1.993673427
0 4 0.773468753 1.748194237 2.970937425 2.478458766
1 5 0.118214112 1.041165238 2.861498509 1.753158127
0 5 0.324957992 1.217722730 3.388696112 1.656588554
0 6 0.056653395 1.361329046 3.430776036 1.592871162
1 6 0.904333005 1.483637306 2.023881029 1.750633756Each pair has two outcomes (disease or not). How do you use R to conduct paired T-test?
I have tried the for loop and the melt function to adjust the format and then the perform paired t test, but it didn't work
你的核心需求是针对每个变量,比较同一配对ID下病例(outcome=1)和对照(outcome=0)的差异,配对t检验的关键是要让每一对的数据精准对应。之前用melt可能方向反了,我们先把数据整理成适合检验的格式,再批量执行检验。
步骤1:准备数据
首先假设你的原始数据框叫df,先统一列名规范(比如把空格替换成下划线,避免后续报错),这里我先模拟一份和你结构一致的数据:
# 加载需要的工具包 library(tidyverse) # 模拟你的配对数据 set.seed(123) df <- tibble( outcome = rep(c(1,0), 6), paired_ID = rep(1:6, each=2), variable_A = rnorm(12, 0.5, 0.2), variable_B = rnorm(12, 1.5, 0.3), variable_C = rnorm(12, 3, 0.4), variable_99 = rnorm(12, 2, 0.2) )
步骤2:数据重塑为宽格式
我们需要把每个变量拆成「病例组」和「对照组」两列,让每一行对应一个配对,这样配对t检验才能识别到对应关系。用pivot_wider实现:
df_wide <- df %>% pivot_wider( id_cols = paired_ID, # 按配对ID分组 names_from = outcome, # 用outcome作为列名后缀 values_from = starts_with("variable") # 所有变量列都拆分 ) # 查看结果:每个变量会变成 variable_X_1(病例)和 variable_X_0(对照) head(df_wide)
步骤3:批量执行配对t检验
现在有两种简洁的方法批量处理所有变量:
方法1:用for循环(新手友好)
# 提取所有变量名(去掉后缀_1) var_names <- str_subset(names(df_wide), "_1$") %>% str_remove("_1$") # 创建空列表存储检验结果 ttest_results <- list() # 循环每个变量执行检验 for (var in var_names) { # 提取病例和对照的数据 case_data <- df_wide[[paste0(var, "_1")]] control_data <- df_wide[[paste0(var, "_0")]] # 执行配对t检验 ttest <- t.test(case_data, control_data, paired = TRUE) # 把结果存入列表,用变量名做标识 ttest_results[[var]] <- ttest } # 查看某个变量的结果,比如variable_A ttest_results[["variable_A"]]
方法2:用purrr包(更高效的tidyverse风格)
如果你熟悉tidyverse,用map函数可以更简洁地批量处理:
# 整理成适合批量处理的格式:每个变量对应一对数据 var_pairs <- var_names %>% set_names() %>% map(~ list(case = df_wide[[paste0(.x, "_1")]], control = df_wide[[paste0(.x, "_0")]])) # 批量执行配对t检验 ttest_results <- var_pairs %>% map(~ t.test(.x$case, .x$control, paired = TRUE)) # 把结果转成数据框,方便查看所有变量的统计量 ttest_summary <- ttest_results %>% map_dfr( ~ tibble( mean_diff = .x$estimate[[1]] - .x$estimate[[2]], t_value = .x$statistic[[1]], p_value = .x$p.value, ci_lower = .x$conf.int[[1]], ci_upper = .x$conf.int[[2]] ), .id = "variable" ) # 查看所有变量的检验结果 print(ttest_summary, n = Inf)
补充:用长格式也能实现
如果你之前习惯用melt转长格式,其实也可以直接在长格式下分组执行检验,核心是确保同一配对ID的两组数据对应:
# 转成变量名-值的长格式 df_long <- df %>% pivot_longer(cols = starts_with("variable"), names_to = "variable", values_to = "value") # 按变量分组执行配对t检验 ttest_results_long <- df_long %>% group_by(variable) %>% summarise( t_test = list(t.test(value ~ outcome, paired = TRUE, data = cur_data())) ) # 查看第一个变量的结果 ttest_results_long$t_test[[1]]
备注:内容来源于stack exchange,提问作者weiri0825

