含假阴性/假阳性的比例置信区间计算问题咨询
Great question—this is a common scenario in diagnostic testing where observed counts don’t directly map to true positives or negatives. Here’s a clear, actionable approach to compute the confidence interval (CI) for the true positive probability π, accounting for known false positive (α) and false negative (β) rates:
Key Background & Simplification
First, let’s clarify the relationship between your observed data and the true underlying probabilities:
- For any individual in your sample, the probability of testing positive is:
p = π*(1 - β) + (1 - π)*α
This combines the chance of a true positive being correctly identified (π*(1-β)) and a true negative being incorrectly labeled positive ((1-π)*α). - Critically, since each individual’s test result is independent, the total number of observed positives
kfollows a Binomial(n, p) distribution. This lets us leverage standard binomial CI methods before transforming toπ.
Step-by-Step Calculation
Assume you have pre-estimated values for α (false positive rate) and β (false negative rate) from a gold-standard validation of your test. If you don’t have these, you’ll need to first validate your test to get these rates—they’re essential for this correction.
1. Compute a CI for the marginal test-positive probability p
Use your preferred binomial CI method on k and n:
- Agresti-Coull: A robust, less conservative alternative to Clopper-Pearson, especially for small samples. Adjust
ktok' = k + 2andnton' = n + 4, then compute the Wald CI:p' = k'/n'SE = sqrt(p'*(1-p')/n')CI_p = [p' - z*SE, p' + z*SE](usez*=1.96for 95% confidence) - Clopper-Pearson: The exact (but conservative) binomial CI. You can calculate this using statistical software or online calculators.
2. Transform the p CI to get the π CI
Since π = (p - α)/(1 - α - β) (let’s call C = 1 - α - β for simplicity), apply this linear transformation to the bounds of your p CI:
π_low = max(0, (p_low - α)/C)π_high = min(1, (p_high - α)/C)
We clamp the values to[0,1]because probabilities can’t fall outside this range.
Example Walkthrough
Let’s say:
- Sample size
n = 100, observed positivesk = 20 - False positive rate
α = 0.05, false negative rateβ = 0.10 C = 1 - 0.05 - 0.10 = 0.85
Step 1: Agresti-Coull CI for p
k' = 20 + 2 = 22,n' = 100 + 4 = 104p' = 22/104 ≈ 0.2115SE = sqrt(0.2115*0.7885/104) ≈ 0.040CI_p = [0.2115 - 1.96*0.040, 0.2115 + 1.96*0.040] ≈ [0.133, 0.290]
Step 2: Transform to π CI
π_low = (0.133 - 0.05)/0.85 ≈ 0.098π_high = (0.290 - 0.05)/0.85 ≈ 0.282- Final 95% CI for true positive probability: [0.10, 0.28]
Handling Uncertainty in α and β
If your estimates of α and β are not precise (e.g., from small validation studies), you’ll need to account for their uncertainty:
- Parametric Bootstrap:
- Resample your validation data to generate new estimates
α*andβ* - Resample your sample data to generate a new observed count
k* - Compute
π* = (k*/n - α*)/(1 - α* - β*) - Repeat thousands of times, then take the 2.5th and 97.5th percentiles of the
π*values as your CI.
- Resample your validation data to generate new estimates
- Delta Method: Calculate the combined variance of
πby accounting for variances inp,α, andβ, then construct a Wald CI. This is more mathematical but can be done with basic variance rules.
Important Notes
- If
C = 1 - α - β = 0, your test has no diagnostic value (the probability of testing positive is the same regardless of true status), so you can’t estimateπ. - Always clamp your final CI to
[0,1]—negative values or values greater than 1 don’t make sense for probabilities.
内容的提问来源于stack exchange,提问作者user3208442

