多重假设检验的样本量确定:Wilcoxon检验与Benjamini–Hochberg校正
Great question—calculating sample sizes for FDR-controlled multiple tests (especially non-parametric ones like Wilcoxon) is trickier than single tests, but there’s a structured way to approach it that works across most test types and FDR methods. Let’s break this down step by step, starting with core concepts and moving to actionable steps and references.
Before diving into calculations, nail down these non-negotiable details—they’ll shape everything else:
- Preset FDR level: e.g., α_FDR = 0.05 (the maximum proportion of false discoveries you’re willing to accept)
- Target statistical power: e.g., 80% or 90% (the chance you’ll detect a true effect when it exists)
- Effect size: For Wilcoxon paired tests, use Cliff’s Delta (a non-parametric measure of distribution overlap; 0.1=small, 0.3=medium, 0.5=large) or a median shift between paired groups
- Number of comparisons: Clarify if you’re only testing treatment groups vs. control, or all pairwise combinations (e.g., 1 control + 3 treatments = 3 comparisons; 4 total groups = 6 pairwise comparisons)
- Pairwise correlation: Since this is a paired design, note how correlated each pair of observations is (e.g., same subject pre/post-test) — this impacts variance estimates.
FDR controls the proportion of false positives among all significant results, unlike single-test Type I error (which controls the chance of any false positive). For sample size, this means you have to account for:
- The total number of tests (m)
- The number of true effects (m₁ = total tests - true null hypotheses)
- The tradeoff between keeping FDR low and maintaining power to detect real effects
Two main approaches work for any test + FDR scenario—simulation is the most reliable for non-parametric tests like Wilcoxon.
A. Simulation-Based Approach (Most Flexible)
Since Wilcoxon tests don’t have simple analytical sample size formulas, simulation is your best bet. Here’s how to implement it:
- Simulate paired data: Generate data matching your expected distribution, effect size, and pairwise correlation. For example, create control observations, then shift treatment observations to reflect your desired effect while preserving paired correlation.
- Run tests + FDR correction: Use
pairwise.wilcox.test()withp.adjust.method = "BH", or manually run individual tests and applyp.adjust(). - Calculate FDR and power:
- FDR = (number of false rejections) / (total rejections) (set to 0 if no rejections)
- Power = (number of true rejections) / (number of true effects)
- Iterate: Adjust sample size until your simulated average FDR stays below your target, and power meets your goal.
Example R Code Snippet
library(tidyverse) # Core parameters alpha_fdr <- 0.05 target_power <- 0.8 n_sim <- 1000 # Number of simulation iterations true_effect_groups <- 3 # Treatments with real effects vs control total_groups <- 1 + true_effect_groups cliffs_delta <- 0.4 # Medium effect size pair_cor <- 0.6 # Correlation between paired observations # Function to simulate data and compute FDR/power simulate_fdr_power <- function(n) { fdr_results <- numeric(n_sim) power_results <- numeric(n_sim) for (i in 1:n_sim) { # Generate paired control data (replace with skewed dist if needed) control <- rnorm(n) # Generate treatment data with effect and paired correlation shift <- qnorm((1 + cliffs_delta)/2) # Convert Cliff's Delta to median shift treatments <- map(1:true_effect_groups, ~ control + shift + rnorm(n, sd = sqrt(1 - pair_cor^2))) # Reshape data for testing df <- tibble(control = control, treatment1 = treatments[[1]], treatment2 = treatments[[2]], treatment3 = treatments[[3]]) %>% pivot_longer(everything(), names_to = "group", values_to = "outcome") # Run pairwise Wilcoxon tests (control vs treatments) test_results <- pairwise.wilcox.test(df$outcome, df$group, paired = TRUE, p.adjust.method = "BH", alternative = "two.sided") p_vals <- test_results$p.value[!is.na(test_results$p.value)] # Classify rejections true_null <- rep(FALSE, length(p_vals)) # All treatments have true effects here rejected <- p_vals < alpha_fdr false_rejects <- sum(rejected & true_null) true_rejects <- sum(rejected & !true_null) total_rejects <- sum(rejected) # Compute metrics fdr <- ifelse(total_rejects == 0, 0, false_rejects / total_rejects) power <- true_rejects / sum(!true_null) fdr_results[i] <- fdr power_results[i] <- power } list(avg_fdr = mean(fdr_results), avg_power = mean(power_results)) } # Test multiple sample sizes n_candidates <- c(20, 30, 40, 50) results <- map_dfr(n_candidates, function(n) { res <- simulate_fdr_power(n) tibble(n = n, avg_fdr = res$avg_fdr, avg_power = res$avg_power) }) print(results)
B. Analytical Approximations (For Parametric or Large Samples)
If your sample size is large enough that Wilcoxon results approximate a normal distribution, you can use this shortcut:
- Calculate the sample size needed for a single test using an adjusted alpha:
α_adjusted = α_FDR * (m / m₁)(a conservative estimate for BH correction). - Use this adjusted alpha in a standard non-parametric sample size calculator (e.g., for Wilcoxon, account for its ~95% efficiency relative to t-tests by multiplying the result by ~1.05).
- Validate with simulation—this is just an approximation.
- Sensitivity analysis: Tweak effect size, pairwise correlation, or FDR level to see how sample size changes. If your effect size is smaller than expected, you’ll need more samples.
- Practical constraints: Balance statistical requirements with real-world limits (e.g., cost, sample availability) — you might need to adjust FDR or power targets if needed.
These steps work regardless of your test type (t-test, chi-square, etc.) or FDR method (BH, BY):
- Lock in core parameters: FDR level, power, effect size, number of tests, and any dependencies between tests.
- Pick simulation or analytics: Simulation for complex/non-parametric designs; analytics for simple parametric tests.
- Iterate and validate: Adjust sample size until your metrics meet targets, then run sensitivity checks to ensure robustness.
- Books:
- Multiple Comparisons and Multiple Tests by Westfall et al. (covers FDR theory and sample size strategies)
- Nonparametric Statistical Methods by Hollander, Wolfe, and Chicken (details on Wilcoxon tests and multiple comparisons)
- R Packages:
pwr: For parametric test sample sizes, adaptable for FDR by adjusting alphafdrtool: Helps with FDR estimation and simulation setupsimstudy: Simplifies generating complex paired/grouped data
内容的提问来源于stack exchange,提问作者František Kaláb

