Stata多重插补数据Pearson相关分析:如何获取P值?
Great question! Your current code correctly pools the Fisher's z-transformed correlations across your imputed datasets, but it misses the critical step of calculating the corresponding p-value. To do this, we need to follow Rubin's rules for multiple imputation pooling, which account for both within-imputation variance (variability within each imputed dataset) and between-imputation variance (differences across imputed datasets).
Here's the full extended code that computes both the pooled correlation and its two-tailed p-value:
mi query local M=10 // Number of imputed datasets // Initialize scalars and matrix to store intermediate results scalar total_z = 0 scalar within_var_sum = 0 matrix z_values = J(`M', 1, .) // Matrix to hold Fisher z-values from each imputation // Loop through each imputed dataset to collect z-values and within-imputation variances mi xeq 1/`M': { correlate v1 v2 local rho = r(rho) local z = atanh(`rho') // Fisher's z-transform matrix z_values[`_mi_m', 1] = `z' scalar total_z = total_z + `z' // Calculate asymptotic variance of z for this dataset: 1/(n-3), where n = sample size local n = e(N) scalar within_var_sum = within_var_sum + (1/(`n' - 3)) } // Compute average Fisher z-value across imputations scalar avg_z = total_z / `M' // Calculate between-imputation variance scalar between_var = 0 forvalues i=1/`M' { scalar between_var = between_var + (z_values[`i', 1] - avg_z)^2 } scalar between_var = between_var / (`M' - 1) // Total variance per Rubin's rules: (within variance / M) + between variance*(1 + 1/M) scalar total_var = (within_var_sum / `M') + between_var*(1 + 1/`M') // Convert average z back to Pearson correlation coefficient scalar pooled_corr = tanh(avg_z) // Calculate z-statistic and two-tailed p-value scalar z_stat = avg_z / sqrt(total_var) scalar p_val = 2 * normal(-abs(z_stat)) // Normal CDF for two-tailed test // Print results with clear formatting di as txt "=== Pooled Pearson Correlation Results ===" di as txt "Pooled correlation coefficient: " as res %0.4f pooled_corr di as txt "Z-statistic: " as res %0.4f z_stat di as txt "Two-tailed p-value: " as res %0.4f p_val
Key Details Explained:
- Fisher's Z-Transform: We still use this to stabilize the variance of the correlation coefficient, which makes pooling across imputations valid.
- Within-Imputation Variance: For each imputed dataset, the asymptotic variance of the Fisher z-value is
1/(n-3)(wherenis the sample size of that dataset). We sum these variances across all imputations. - Between-Imputation Variance: This measures how much the z-values vary across different imputed datasets—critical for accounting for uncertainty from missing data.
- Rubin's Pooling: The total variance combines both within and between components, adjusted for the number of imputations (
M). - P-Value Calculation: We use the standard normal distribution to compute the two-tailed p-value from the z-statistic (average z divided by the square root of total variance).
Note:
If all your imputed datasets have the same sample size (which is typical for multiple imputation), you can simplify the within_var_sum calculation to M*(1/(n-3)) (replace n with your sample size) to save a bit of computation—though the code above works even if sample sizes differ.
内容的提问来源于stack exchange,提问作者ksroogl

