如何对R数据框中各tax分组和指定列批量执行fisher.test检验
需求说明
我们需要对给定数据框按以下规则执行批量统计检验:
- 分组维度:按
tax字段的每个分组单独计算 - 检验列:对每个tax分组下的目标列(示例中为ABCB1、ABL1)分别执行
fisher.test()检验 - 列联表构造规则:根据现有数据行计算生成,示例列联表格式如下:
| ABCB1 | NotABCB1 | |
|---|---|---|
| tax1Present | 42 | 1 |
| tax1NotPresent | 20 | 3 |
计算逻辑:
- 42 = 总Present数 - tax1Present行对应ABCB1列的值(即43-1)
- 1 = tax1Present行对应ABCB1列的单元格值
- 20 = 总NotPresent数 - tax1NotPresent行对应ABCB1列的值(即23-3)
- 3 = tax1NotPresent行对应ABCB1列的单元格值
所用数据集
数据集构造代码:
df <- structure(list(group = c("tax1Present", "tax1NotPresent", "tax2Present", "tax2NotPresent", "tax3Present", "tax3NotPresent", "tax4Present", "tax4NotPresent", "tax5Present", "tax5NotPresent"), ABCB1 = c(1L, 3L, 4L, 5L, 3L, 6L, 6L, 3, 2, 6L), ABL1 = c(18L, 14, 12L, 9L, 1L, 5L, 0L, 0L, 7L, 0L), Present = c(43L, 43, 23L, 23, 9L, 9, 7L, 7, 20, 20L),NotPresent = c(23, 23, 18, 18, 7L, 7L, 10L, 10L, 10, 10L), tax = c("tax1", "tax1", "tax2", "tax2", "tax3", "tax3", "tax4", "tax4", "tax5", "tax5")), row.names = c(NA, 10L), class = "data.frame")
数据集预览:
> df group ABCB1 ABL1 Present NotPresent tax 1 tax1Present 1 18 43 23 tax1 2 tax1NotPresent 3 14 43 23 tax1 3 tax2Present 4 12 23 18 tax2 4 tax2NotPresent 5 9 23 18 tax2 5 tax3Present 3 1 9 7 tax3 6 tax3NotPresent 6 5 9 7 tax3 7 tax4Present 6 0 7 10 tax4 8 tax4NotPresent 3 0 7 10 tax4 9 tax5Present 2 7 20 10 tax5 10 tax5NotPresent 6 0 20 10 tax5
实现代码
结合dplyr分组处理和purrr批量迭代即可完成计算,代码如下:
library(tidyverse) # 指定要检验的目标列 target_cols <- c("ABCB1", "ABL1") # 按tax分组批量计算 result <- df %>% group_by(tax) %>% group_modify(~{ # 提取当前分组的两行数据 present_row <- filter(.x, str_detect(group, "Present$")) not_present_row <- filter(.x, str_detect(group, "NotPresent$")) total_p <- present_row$Present total_np <- present_row$NotPresent # 对每个目标列计算fisher检验 map_dfr(target_cols, function(col){ # 构造列联表 cont_table <- matrix( c(total_p - present_row[[col]], present_row[[col]], total_np - not_present_row[[col]], not_present_row[[col]]), nrow = 2, byrow = TRUE ) # 执行fisher检验 ft <- fisher.test(cont_table) # 返回结果 tibble( target_col = col, p_value = ft$p.value, odds_ratio = as.numeric(ft$estimate), conf_low = ft$conf.int[1], conf_high = ft$conf.int[2] ) }) }) %>% ungroup() # 查看结果 print(result)
运行后会输出每个tax分组、每个目标列对应的Fisher精确检验结果,包含p值、优势比及95%置信区间,可直接根据阈值筛选显著结果。
内容的提问来源于stack exchange,提问作者user2300940
相关产品推荐
相关产品推荐

