You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中fisher.test两次计算结果不同的原因咨询

为何两次Fisher精确检验结果不一致?

我原本以为在R中两次运行相同的fisher.test()会得到一致结果,但实际却不一样,操作过程如下:

首先从数据生成列联表:

> SAtab1 <- table(dt2$ACE_SA_BIN, dt2$IHD_other_healthy)
> SAtab1
   
      0   1   2
  0 345 577  29
  1  10  25   2

随后运行Fisher检验:

> fisher.test(SAtab1, workspace = 2e8, simulate.p.value=TRUE)

    Fisher's Exact Test for Count Data with simulated p-value (based on 2000 replicates)

data:  SAtab1
p-value = 0.2969
alternative hypothesis: two.sided

之后我尝试传入矩阵而非列联表,创建矩阵:

> SAtab2 <- 
+   matrix(c(345,10,577,25,29,2),
+          nrow = 2,
+          dimnames = list(Gender = c("no", "yes"),
+                          Illness =c("no illness","other than IHD", "HD")))
> SAtab2
      Illness
Gender no illness other than IHD HD
   no         345            577 29
   yes         10             25  2

再次运行Fisher检验:

> fisher.test(SAtab2, workspace = 2e8, simulate.p.value=TRUE)

    Fisher's Exact Test for Count Data with simulated p-value (based on 2000 replicates)

data:  SAtab2
p-value = 0.2829
alternative hypothesis: two.sided

差异原因

两次结果的微小差异完全来自于你设置了simulate.p.value=TRUE参数:

  • 当开启模拟p值时,R会通过随机抽样生成2000个重复样本(默认replicates=2000)来估算p值,每次抽样的随机数种子不同,得到的模拟样本集合就有差异,最终计算出的p值自然会有小幅波动。
  • 你传入的SAtab1和SAtab2本质是完全相同的频数分布,只是数据类型(table vs matrix)和维度名称不同,这对Fisher检验的计算逻辑没有影响。

如果想要得到完全一致的结果,可以在每次运行检验前设置相同的随机数种子,比如:

set.seed(123)
fisher.test(SAtab1, workspace = 2e8, simulate.p.value=TRUE)

set.seed(123)
fisher.test(SAtab2, workspace = 2e8, simulate.p.value=TRUE)

这样两次模拟会使用相同的随机序列,得到的p值就会完全一致。

内容的提问来源于stack exchange,提问作者Nneka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 15:08:01