如何向量化R语言fisher.test函数以批量分析多变量?
批量执行Fisher精确检验并提取P值(解决group_modify报错)
问题场景
需要批量对多个特征变量与结局变量执行Fisher精确检验,最终生成包含变量名和对应P值的结果表。但使用group_modify尝试批量处理时,触发错误:
no method 'count' applicable for an object of class c("integer", "numeric")
示例数据集
library(tidyverse) library(broom) set.seed(123) # 设置随机种子确保结果可复现 n <- 20 outcome <- rbinom(n, size = 1, prob = 0.7) feature1 <- rbinom(n, size = 2, prob = 0.5) + 10 feature2 <- rbinom(n, size = 1, prob = 0.5) + 20 df <- tibble(outcome, feature1, feature2)
单变量检验参考代码(正确逻辑)
df %>% select(outcome, feature1) %>% count(outcome, feature1) %>% pivot_wider(names_from = feature1, values_from = n) %>% fisher.test() %>% tidy()
错误原因
原批量代码存在两处逻辑问题:
group_modify(~count(outcome, value))未针对分组后的子数据框操作,直接调用全局变量导致类型不匹配,触发报错;- 分组后未在每个组内完成完整的检验流程,反而直接对整体数据框执行
fisher.test,不符合批量处理的逻辑。
修正后的批量处理方案
方案一:基于group_modify的正确实现
df %>% pivot_longer(c("feature1", "feature2"), names_to = "variable", values_to = "value") %>% group_by(variable) %>% group_modify(function(.x, .y) { # 对每个分组的子数据框执行完整检验流程 .x %>% count(outcome, value) %>% pivot_wider(names_from = value, values_from = n, values_fill = 0) %>% # 补0避免列缺失导致检验失败 select(-outcome) %>% as.matrix() %>% # fisher.test需要矩阵输入 fisher.test() %>% tidy() %>% mutate(variable = .y$variable) # 绑定当前分组的变量名 }) %>% select(variable, p.value) %>% ungroup()
方案二:基于purrr::map的实现
如果更习惯用迭代式写法,可采用map_dfr批量处理每个变量:
feature_names <- c("feature1", "feature2") map_dfr(feature_names, function(var) { df %>% select(outcome, all_of(var)) %>% count(outcome, !!sym(var)) %>% pivot_wider(names_from = !!sym(var), values_from = n, values_fill = 0) %>% select(-outcome) %>% as.matrix() %>% fisher.test() %>% tidy() %>% mutate(variable = var) }) %>% select(variable, p.value)
输出结果
执行上述代码后,会得到符合预期的结果表:
# A tibble: 2 × 2 variable p.value <chr> <dbl> 1 feature1 0.805 2 feature2 0.582
内容的提问来源于stack exchange,提问作者T. Walter
相关产品推荐
相关产品推荐

