scipy.stats.ttest_ind结果困惑:不同参数下结果差异原因解析
scipy.stats.ttest_ind Behavior: Pooled vs Welch's T-Test First, let's clarify the core difference between the two modes of ttest_ind:
equal_var=True: Uses the pooled variance t-test, which assumes the two populations have equal variances and combines sample variances into a single pooled estimate.equal_var=False: Uses Welch's t-test, which does not assume equal variances and calculates separate variance estimates for each group, weighted by sample size.
1. Same Sample Size, Different Variances: Minimal Result Differences
When your two samples are the same size (n1 = n2), here's why the two tests produce nearly identical results:
Key Formula Comparison
- Pooled standard error:
For equal sample sizes, the pooled variance simplifies to the average of the two sample variances (adjustments for degrees of freedom are negligible with large samples). The standard error becomes:SE_pooled = sqrt( (s1² + s2²)/n ) - Welch's standard error:
For equal sample sizes, this also simplifies to:
The standard error is exactly the same for both tests.SE_welch = sqrt( s1²/n + s2²/n ) = sqrt( (s1² + s2²)/n )
Why P-Values Are Almost Identical
Since the standard error matches, the t-statistic is identical. The only difference is degrees of freedom:
- Pooled test uses
n1 + n2 - 2(998 in your example). - Welch's test uses the Satterthwaite approximation, which reduces df (730 in your example).
With large sample sizes (500 each), the t-distribution is nearly identical to the normal distribution. A small difference in df has almost no impact on the p-value, hence the minimal variation you see.
2. Different Sample Sizes + Different Variances: Large Result Differences
When sample sizes and variances both differ, the two tests diverge significantly because:
Pooled Variance Biases Towards the Larger Sample
The pooled variance gives more weight to the sample with more data points. In your example:
rvs1: 500 samples, variance ~100 (scale=10)rvs4: 100 samples, variance ~400 (scale=20)
The pooled variance is dominated by the larger, lower-variance sample:
s_p² = [(500-1)*100 + (100-1)*400]/(500+100-2) ≈ 149.66
This lowers the overall standard error, leading to a larger t-statistic (-5.125) and a much smaller p-value (4.0e-07).
Welch's Test Treats Variances Separately
Welch's test doesn't pool variances—it calculates the standard error as the sum of each variance divided by its sample size:
SE_welch = sqrt(100/500 + 400/100) = sqrt(0.2 +4) ≈2.05
Here, the smaller, higher-variance sample contributes far more to the standard error (4 vs 0.2). This results in a larger standard error, a smaller t-statistic (-3.577), and a higher p-value (0.00051).
Degrees of Freedom Amplify the Difference
Welch's test uses a much smaller degrees of freedom (111.951 vs 598 for pooled). A smaller df makes the t-distribution wider, which further increases the p-value compared to the pooled test (which uses a narrower, normal-like distribution).
Summary
- Same sample size: Pooled and Welch's tests produce nearly identical results because their standard errors are equal, and large df minimizes p-value differences.
- Different sizes + variances: Pooled test biases towards the larger sample's variance, leading to smaller SE, larger t-statistic, and smaller p-value. Welch's test accounts for unequal variances separately, resulting in larger SE, smaller t-statistic, and larger p-value.
内容的提问来源于stack exchange,提问作者Dan Bolfter

