如何用统计显著性检验验证预测值与实际值样本的相似性?
Great question—you’re spot on that a non-significant paired t-test doesn’t equate to similarity between your samples. That test only tells us we can’t reject the null hypothesis of equal means; it gives no insight into how closely individual predicted values match their actual counterparts. To statistically validate similarity (often referred to as equivalence testing), here are the most reliable methods for paired data:
1. Paired Equivalence t-Test (Two One-Sided Tests, TOST)
This is the gold standard for equivalence testing. Here’s how it works:
- First, define a practical equivalence interval: decide on a threshold Δ where you’d consider a difference between predicted and actual values (
|predicted - actual| ≤ Δ) to be "similar enough" for your use case (e.g., Δ = 5% of the actual value range). - Run two one-sided t-tests:
- Test that the mean difference is greater than -Δ (rejecting the null that mean difference ≤ -Δ)
- Test that the mean difference is less than Δ (rejecting the null that mean difference ≥ Δ)
- If both tests are statistically significant, you can conclude your paired samples are equivalent (similar) within your defined threshold.
2. Bland-Altman Analysis (With Statistical Validation)
While the Bland-Altman plot is a visualization tool, you can pair it with statistical checks to confirm similarity:
- Calculate the mean difference between paired values and its 95% confidence interval (CI).
- If this entire 95% CI falls inside your pre-defined equivalence interval, it provides statistical evidence that the two samples are sufficiently similar.
- Bonus: You can also test if the difference varies with the magnitude of the actual values (e.g., via a correlation test between the mean of paired values and their difference) to ensure consistent similarity across the range.
3. Intraclass Correlation Coefficient (ICC)
ICC measures the consistency (agreement) between paired measurements, rather than just comparing means:
- An ICC value close to 1 indicates high consistency between predicted and actual values (individual pairs track closely together).
- Make sure to choose the correct ICC model based on your study design (e.g., a two-way random effects model if both predicted and actual values are considered random samples from a population).
- Unlike Pearson correlation, ICC accounts for both the correlation between pairs and the similarity of their means, making it better suited for assessing overall similarity.
4. Lin's Concordance Correlation Coefficient
This metric combines two key aspects of similarity:
- The correlation between predicted and actual values (precision)
- The deviation of the mean difference from zero (accuracy)
- A value of 1 means perfect concordance, while values closer to 0 indicate poor similarity. You can test if the coefficient is significantly greater than a threshold you define (e.g., 0.9) to validate similarity.
Key Notes Before You Start
- Define "similarity" first: All these methods rely on a pre-specified, practically meaningful threshold (equivalence interval or concordance cutoff). Statistical tests can’t tell you what’s "similar"—you have to define that based on your field or use case.
- Sample size matters: Equivalence tests require larger sample sizes than standard t-tests to detect similarity. A small sample might lead to non-significant results even if the samples are similar.
内容的提问来源于stack exchange,提问作者Itai Sevitt

