如何为数据框每行生成列联表并批量执行Fisher精确检验?
这里有两种高效的批量处理方法,分别用purrr(函数式编程风格)和dplyr(管道流风格)来实现你的需求,同时帮你处理边界情况(比如物种计数没有下降的场景):
1. 先准备模拟数据
df <- data.frame( Species = c("cat", "dog", "bird"), `2016` = c(14, 16, 10), `2017` = c(8, 12, 5), stringsAsFactors = FALSE )
2. 方法一:用purrr批量处理(函数式风格)
先写一个处理单个物种的函数,它会自动生成指定格式的列联表,并执行Fisher精确检验:
library(purrr) process_single_species <- function(species, count_2016, count_2017) { # 计算2017年的absent数量,避免出现负数 absent_2017 <- max(count_2016 - count_2017, 0) # 构建你需要的列联表结构 contingency_table <- matrix( c(count_2016, count_2017, 0, absent_2017), nrow = 2, dimnames = list( Status = c("present", "absent"), Year = c("2016", "2017") ) ) # 只有当计数确实下降时才执行检验,否则返回提示 if (absent_2017 == 0) { warning(glue::glue("物种 {species} 没有计数下降,跳过Fisher检验")) fisher_result <- NULL } else { # 指定备择假设为"greater",对应检验2017年present的比例显著更低(即计数下降显著) fisher_result <- fisher.test(contingency_table, alternative = "greater") } # 返回包含所有信息的列表 list( Species = species, Contingency_Table = contingency_table, Fisher_Test_Result = fisher_result ) }
然后用pmap批量应用到每一行数据:
# 生成所有物种的结果 all_results <- pmap(df, process_single_species) # 查看单个物种的结果,比如cat all_results[[1]]$Contingency_Table all_results[[1]]$Fisher_Test_Result
3. 方法二:用dplyr管道流处理(更适合数据框操作)
如果你习惯用dplyr的管道语法,可以这样实现,结果会直接整合在数据框中:
library(dplyr) library(purrr) df_results <- df %>% rowwise() %>% mutate( # 计算absent_2017并构建列联表 absent_2017 = max(`2016` - `2017`, 0), contingency_table = list( matrix( c(`2016`, `2017`, 0, absent_2017), nrow = 2, dimnames = list(c("present", "absent"), c("2016", "2017")) ) ), # 执行Fisher检验,跳过无下降的物种 fisher_test = list( if (absent_2017 > 0) { fisher.test(contingency_table, alternative = "greater") } else { NULL } ) ) %>% ungroup() %>% # 提取检验的关键结果(p值、优势比)到单独列,方便查看 mutate( p_value = map_dbl(fisher_test, ~ if(!is.null(.)) .$p.value else NA_real_), odds_ratio = map_dbl(fisher_test, ~ if(!is.null(.)) .$estimate else NA_real_) ) # 查看整理后的结果 df_results %>% select(Species, p_value, odds_ratio, contingency_table)
关键说明
- 设置
alternative = "greater"是为了匹配你的需求:检验计数是否显著下降,这个备择假设对应"2017年present的比例显著低于2016年"的场景。 - 加入
max(count_2016 - count_2017, 0)的处理,避免当2017年计数高于2016年时出现负数的absent值,同时自动跳过这类无下降物种的检验。 - 两种方法都能批量生成你需要的列联表,并完成Fisher检验,结果可以方便地提取和后续分析。
内容的提问来源于stack exchange,提问作者KNN
相关产品推荐
相关产品推荐

