R语言power.prop.test函数参数不可互换的原因问询(A/B测试场景)
Great question! This is a common gotcha when working with power calculations for rare-event A/B tests in R, and the discrepancy boils down to two key factors:
1. You’re solving for different unknowns in each run
First, remember that R's power.prop.test requires exactly one parameter to be NULL (the value you want to calculate) while the rest are fixed. Your two runs are fundamentally different calculations:
- First run: You likely fixed your baseline conversion rate
p1, desired power (e.g., 0.8), significance level (sig.level, e.g., 0.05), and either solved forn(sample size) while gettingp2as the minimum detectable effect (MDE) for thatn, or fixednand solved for thep2that would be detectable at that sample size. - Second run: You took that derived
p2, fixedp1, power, andsig.level, then solved for thenneeded to detect that specificp2vsp1.
In a perfect theoretical scenario, these should align—but real-world approximation errors break this consistency, especially with your data.
2. Normal approximation fails for extremely small proportions
Your p1 and p2 are tiny (~0.00015 and ~0.000187), which qualify as rare events. The power.prop.test function relies on the normal approximation to the binomial distribution to compute power and sample sizes. This approximation works well when proportions are moderate and sample sizes are large enough, but it breaks down dramatically when probabilities are close to 0 (or 1).
The approximation error introduces small biases in the first calculation of p2, and when you feed that biased p2 back into the function to compute n, the error compounds—resulting in the drastically smaller sample size you see.
What to do instead
For rare-event A/B tests (like yours with sub-0.1% conversion rates):
- Avoid the normal approximation entirely. Use methods based on the Poisson distribution (which is better suited for counting rare events) or exact binomial power calculations.
- Double-check that all parameters (like
alternativefor one-sided/two-sided tests,power, andsig.level) are identical across runs—even a small change here can shift results significantly.
内容的提问来源于stack exchange,提问作者Nick Knauer

