You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否采用Exact Fisher Test对比两款iPhone医学数据检测算法的结果?

关于Fisher精确检验在你的算法对比中的适用性分析

Hey there! Let's break down your questions step by step—you've got two key points here: whether Fisher's exact test is the right fit for your 3-category results, and what that p-value <0.05 actually tells you about algorithm performance.

1. Is Fisher's Exact Test appropriate here?

First off, standard Fisher's exact test is designed for 2×2 contingency tables—it checks if two binary variables are independent. But your results have three categories: SUCCESS, FAIL, INVALID, so your data is a 2×3 table (2 algorithms × 3 result types).

That means using a basic 2×2 Fisher test here is not the right call. Instead, you have two better options:

  • Pearson's Chi-Squared Test: If most of your expected cell counts are ≥5 (with 56 samples, this is probably true unless some categories are super rare), this is the go-to for testing if the distribution of results differs between algorithms.
  • Fisher-Freeman-Halton Test: This is the extended version of Fisher's exact test for tables larger than 2×2. It’s ideal if you have small expected counts (though 56 samples should be manageable). Good news: if you passed a 2×3 table to R’s fisher.test() function, it automatically uses this extended version—so if that’s what you did, you’re on the right track! Just make sure you didn’t incorrectly collapse your categories into a 2×2 table first.

2. What does p<0.05 actually mean?

A small p-value tells you that the result distributions of the two algorithms are statistically significantly different—but it doesn’t directly mean one algorithm is "better." "Better" depends on your business definition of success:

  • Do you care most about maximizing SUCCESS rates? Minimizing FAILs? Reducing INVALIDs? Or maybe focusing on SUCCESS rate among valid (non-INVALID) samples?
  • For example: Suppose Algorithm A has 20 SUCCESS, 5 FAIL, 3 INVALID; Algorithm B has 15 SUCCESS, 4 FAIL, 9 INVALID. The overall distribution differs, but the gap is driven by INVALIDs. You need to decide if more INVALIDs make B worse, or if your INVALID definition needs tweaking.

So here’s what to do next:

  • Define what "better" means for your use case. Let’s say it’s "highest SUCCESS rate among valid samples."
  • Filter out INVALID entries, create a 2×2 table (SUCCESS/FAIL vs. Algorithm A/B), then run the standard Fisher’s exact test on that—this will directly test if one algorithm has a significantly higher success rate in valid cases.
  • Calculate an effect size like Cramér's V for your original 2×3 table. P-values only tell you if a difference exists; effect sizes tell you how big that difference is (and whether it matters in practice). Cramér's V ranges from 0 (no association) to 1 (perfect association).

3. Quick R Code Examples

Let’s assume your data is in a data frame df with columns algorithm (values "A" and "B") and result (values "SUCCESS", "FAIL", "INVALID"):

Generate the contingency table

result_table <- table(df$algorithm, df$result)
print(result_table)

Run the appropriate tests

# Extended Fisher's test for 2×3 table
fisher.test(result_table)

# Chi-squared test (if expected counts are sufficient)
chisq.test(result_table)

Focus on valid samples (exclude INVALID)

valid_samples <- df[df$result != "INVALID", ]
valid_table <- table(valid_samples$algorithm, valid_samples$result)
fisher.test(valid_table)

Calculate Cramér's V for effect size

library(rcompanion)
cramerV(result_table)

内容的提问来源于stack exchange,提问作者Fab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:49:24