Stata面板数据多插补后取均值替代合并结果是否可行?
Great question—this is a common point of confusion when working with multiple imputation (MI) in Stata, so let’s break down why averaging imputed values isn’t a reliable substitute for the standard MI workflow.
Why Averaging Imputed Values Is Not Recommended for Statistical Inference
- You’ll erase critical uncertainty: The 10 different imputed values you generated exist to capture the uncertainty around what the true missing values might be. Averaging these values flattens that variation entirely. When you run analyses on this single "average-imputed" dataset, your standard errors will be too small, confidence intervals too narrow, and p-values too low—leading you to overstate the significance of your results.
- It violates Rubin’s MI framework: The gold standard for MI (Rubin’s rules) explicitly accounts for two sources of variance:
- Within-imputation variance: Variation in results from a single imputed dataset
- Between-imputation variance: Variation across all 10 imputed datasets (this is the uncertainty from missing data)
Averaging imputed values only uses the within-imputation variation, completely ignoring the between-imputation component. This makes your statistical results statistically invalid.
What You Should Do Instead
Stata makes it easy to follow the correct MI workflow with the mi estimate command. After running your 10 imputations, just run your analysis within the mi estimate wrapper, and Stata will automatically apply Rubin’s rules to merge results across all imputed datasets.
For example, if you’re running a regression:
mi estimate, cmdreg: regress depvar indepvar1 indepvar2 indepvar3
This command will output combined coefficients, adjusted standard errors, and valid p-values that properly account for missing data uncertainty.
When Might Averaging Be Okay?
If you only need a single dataset for descriptive purposes (like creating summary tables or basic visualizations), averaging imputed values can be a quick simplification. But you must clearly note that this dataset does not reflect the uncertainty of the missing data, and it should never be used for formal hypothesis testing or regression analysis.
内容的提问来源于stack exchange,提问作者Badalyan

