You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为数据框每行生成列联表并批量执行Fisher精确检验?

这里有两种高效的批量处理方法,分别用purrr(函数式编程风格)和dplyr(管道流风格)来实现你的需求,同时帮你处理边界情况(比如物种计数没有下降的场景):

1. 先准备模拟数据

df <- data.frame(
  Species = c("cat", "dog", "bird"),
  `2016` = c(14, 16, 10),
  `2017` = c(8, 12, 5),
  stringsAsFactors = FALSE
)

2. 方法一:用purrr批量处理(函数式风格)

先写一个处理单个物种的函数,它会自动生成指定格式的列联表,并执行Fisher精确检验:

library(purrr)

process_single_species <- function(species, count_2016, count_2017) {
  # 计算2017年的absent数量,避免出现负数
  absent_2017 <- max(count_2016 - count_2017, 0)
  
  # 构建你需要的列联表结构
  contingency_table <- matrix(
    c(count_2016, count_2017,
      0, absent_2017),
    nrow = 2,
    dimnames = list(
      Status = c("present", "absent"),
      Year = c("2016", "2017")
    )
  )
  
  # 只有当计数确实下降时才执行检验,否则返回提示
  if (absent_2017 == 0) {
    warning(glue::glue("物种 {species} 没有计数下降,跳过Fisher检验"))
    fisher_result <- NULL
  } else {
    # 指定备择假设为"greater",对应检验2017年present的比例显著更低(即计数下降显著)
    fisher_result <- fisher.test(contingency_table, alternative = "greater")
  }
  
  # 返回包含所有信息的列表
  list(
    Species = species,
    Contingency_Table = contingency_table,
    Fisher_Test_Result = fisher_result
  )
}

然后用pmap批量应用到每一行数据:

# 生成所有物种的结果
all_results <- pmap(df, process_single_species)

# 查看单个物种的结果,比如cat
all_results[[1]]$Contingency_Table
all_results[[1]]$Fisher_Test_Result

3. 方法二:用dplyr管道流处理(更适合数据框操作)

如果你习惯用dplyr的管道语法,可以这样实现,结果会直接整合在数据框中:

library(dplyr)
library(purrr)

df_results <- df %>%
  rowwise() %>%
  mutate(
    # 计算absent_2017并构建列联表
    absent_2017 = max(`2016` - `2017`, 0),
    contingency_table = list(
      matrix(
        c(`2016`, `2017`, 0, absent_2017),
        nrow = 2,
        dimnames = list(c("present", "absent"), c("2016", "2017"))
      )
    ),
    # 执行Fisher检验,跳过无下降的物种
    fisher_test = list(
      if (absent_2017 > 0) {
        fisher.test(contingency_table, alternative = "greater")
      } else {
        NULL
      }
    )
  ) %>%
  ungroup() %>%
  # 提取检验的关键结果(p值、优势比)到单独列,方便查看
  mutate(
    p_value = map_dbl(fisher_test, ~ if(!is.null(.)) .$p.value else NA_real_),
    odds_ratio = map_dbl(fisher_test, ~ if(!is.null(.)) .$estimate else NA_real_)
  )

# 查看整理后的结果
df_results %>%
  select(Species, p_value, odds_ratio, contingency_table)

关键说明

  • 设置alternative = "greater"是为了匹配你的需求:检验计数是否显著下降,这个备择假设对应"2017年present的比例显著低于2016年"的场景。
  • 加入max(count_2016 - count_2017, 0)的处理,避免当2017年计数高于2016年时出现负数的absent值,同时自动跳过这类无下降物种的检验。
  • 两种方法都能批量生成你需要的列联表,并完成Fisher检验,结果可以方便地提取和后续分析。

内容的提问来源于stack exchange,提问作者KNN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:53:50