Stata中含unit specific trends的Dif-in-Dif回归adjusted R²过高问题咨询
Hey there, let’s unpack why you’re seeing those near-perfect adjusted R² values (even 0.99!) in your DID model with unit-specific trends, and walk through how to verify if this is a problem or just a quirk of your setup.
First, why the sky-high R²?
This is actually pretty common in panel models with unit-specific trends, especially given your dataset structure:
- Unit-specific trends are powerful fitters: You’re adding 201 separate linear trends (one per country) to your model. With 36 years of data per country, these trends will soak up almost all the long-term, country-specific variation in your dependent variable—think things like gradual GDP growth, population increases, or policy trajectories unique to each nation. If your DV has strong time-dependent patterns at the country level, these trends will explain most of its variation.
- Parameter count relative to observations: You’ve got ~450 parameters for 5000 observations. While that’s not an extreme ratio (roughly 11 observations per parameter), combining country fixed effects, year fixed effects, 5 controls, and 201 trend terms means your model is accounting for a huge portion of the data’s structure.
- DV type matters: If your dependent variable is a macroeconomic or slow-changing variable (e.g., GDP per capita, life expectancy), it naturally has strong linear trends at the country level. The unit-specific trends will essentially "fit" those trends almost perfectly, driving R² way up.
What to do next to validate your model
Don’t panic—high R² doesn’t automatically mean your model is wrong, but you should check a few key things to ensure your core DID estimate is reliable:
Test the model without unit-specific trends:
Run a simpler DID specification first to compare results:xtset country year reg y treat i.year controls, robust feThen compare it to your full model with trends:
reg y treat i.year controls i.country#c.year, robust feLook at how your treatment effect coefficient changes—does it stay significant? Does its magnitude make economic sense? If the treatment effect disappears or becomes implausible when adding trends, that’s a red flag. Also, note the drop in R² when removing trends—this will tell you how much variation the trends are actually explaining.
Don’t fixate on adjusted R² for DID models:
Adjusted R² measures overall fit, but in DID, we care far more about the causal effect of the treatment, not how well the model predicts the DV. Even if R² is 0.99, if your treatment effect is statistically significant and aligns with your theoretical expectations, the model might still be valid.Decompose your dependent variable’s variation:
Use Stata’sxtsumcommand (afterxtset-ing your data) to split your DV into between-country variation and within-country over-time variation:xtsum yUnit-specific trends explain within-country variation. If most of your DV’s variation is within-country (high
withinpercentage), that’s why the trends are driving R² up—and that’s totally normal for many macro panel datasets.Check for overfitting (if you’re worried):
Split your dataset into a training and test sample (e.g., use the first 30 years to estimate the model, then predict the last 6 years). If the predicted values are close to the actual values, the high R² is just a reflection of your data’s structure, not overfitting. If the predictions are way off, you might need to revisit your model specification.Verify no data anomalies:
Quick checks to rule out silly mistakes:- Are there any extreme outliers in your DV? Use
summarize y, detailorhistogram yto check. - Is there accidental multicollinearity? While fixed effects models mitigate this, you can run
estat vifon a model without fixed effects to see if controls are highly correlated (though take VIF results with a grain of salt for panel models).
- Are there any extreme outliers in your DV? Use
Final takeaway
In most cases, a high adjusted R² in a DID model with unit-specific trends is not an error—it’s just the model doing its job of accounting for country-specific time patterns. The real test is whether your treatment effect is robust, statistically significant, and makes sense for your research question.
内容的提问来源于stack exchange,提问作者Angelo Cozzubo

