SEM拟合优度与模型修正:混合方法论文中SEM构建技术问询
Hey Janina, great question—combining qualitative deep interview insights with SEM for theory building is a really rigorous, mixed-methods approach, so it makes total sense to get clear on how to assess model fit and refine your structure appropriately. Let’s walk through this tailored to your specific model: a binary observed predictor, one latent outcome (3 ordered items), and one latent variable (7 ordered items).
Since your model includes ordered scale items (not continuous variables) and a binary observed predictor, standard fit indices need targeted adjustments, and you’ll want to balance statistical fit with your theoretical framework.
1.1 Core Fit Indices to Prioritize
Focus on a mix of absolute, incremental, and parsimonious indices to get a holistic view:
- Absolute Fit:
- RMSEA (Root Mean Square Error of Approximation): Aim for < 0.06 for excellent fit; < 0.08 is acceptable. For ordered data, use the robust weighted least squares (WLSMV) estimator (standard in tools like
Mplusorlavaanfor R)—ML estimators aren’t ideal for categorical/ordered items. - SRMR (Standardized Root Mean Square Residual): Target < 0.08; this measures residual covariance between items, which is critical for capturing relationships in ordered scales.
- RMSEA (Root Mean Square Error of Approximation): Aim for < 0.06 for excellent fit; < 0.08 is acceptable. For ordered data, use the robust weighted least squares (WLSMV) estimator (standard in tools like
- Incremental Fit:
- CFI (Comparative Fit Index) & TLI (Tucker-Lewis Index): Look for values > 0.95 for top-tier fit; > 0.90 is still acceptable. These compare your model to a baseline null model, which is perfect for theory-building where you’re testing a targeted structure.
- Parsimonious Fit:
- Chi-Square/df Ratio: Aim for a 2-3:1 ratio (note: raw chi-square is sample-size sensitive, so don’t rely on it alone).
- AIC (Akaike Information Criterion) & BIC (Bayesian Information Criterion): Lower values mean better fit relative to model complexity—use these if you’re comparing nested versions of your model during refinement.
1.2 Special Considerations for Your Model Components
- Ordered Latent Variables: Ensure your estimator is set to handle categorical data. For example, in
lavaan, you’d useestimator = "WLSMV"andordered = c(...)to tag your scale items. Avoid indices designed for continuous data like GFI here. - Binary Predictor: This non-latent variable doesn’t need special fit adjustments, but double-check that its direct path to the latent variables is statistically significant (p < 0.05) and aligns with your qualitative interview insights.
Since this is a theory-building mixed-methods study, corrections shouldn’t just be statistically driven—they need to tie back to the themes and insights from your deep interviews. Here’s a structured approach:
2.1 Statistically Driven Initial Refinement
Start with statistical signals, but always ground changes in theory:
- Tweak the Measurement Model First:
- Check item loadings: If any item in your 3-item or 7-item latent variable has a loading < 0.5, consider removing it—but only if it doesn’t contradict your qualitative findings (e.g., if the item never came up as a core construct component in interviews).
- Address residual covariances: If SRMR is high, look for item pairs with large residual covariances. If those items are conceptually linked (e.g., they tap the same sub-dimension identified in interviews), you can add a residual covariance path between them.
- Adjust Structural Paths:
- If the direct path from your binary predictor to the latent outcome is non-significant, revisit your interview data—did participants describe an indirect relationship? Your 7-item latent variable might act as a mediator (binary predictor → 7-item latent → 3-item latent outcome), which you can test by adding that indirect path.
2.2 Theory-Driven Refinement (Leverage Your Qualitative Insights)
This is where your mixed-methods design adds unique value—use interview data to guide corrections that make theoretical sense:
- Add/Remove Constructs: If interviews consistently highlighted a factor you didn’t include initially, consider adding a latent variable (e.g., a moderator that interacts with your binary predictor). Conversely, if a latent variable’s items don’t align with interview themes, refine or remove the item set.
- Redefine Path Relationships: If interviews showed the binary predictor first impacts the 7-item latent variable, which then affects the 3-item outcome, adjust your model to reflect this mediating relationship instead of a direct path.
- Refine Item Operationalization: If an item has a low loading, check if its phrasing doesn’t match how participants talked about the construct. Reword the item (if you can re-collect data) or justify removing it using qualitative evidence.
2.3 Avoid Over-Correction Pitfalls
- Don’t chase perfect fit at the cost of theoretical validity: A model with slightly higher RMSEA but strong alignment with interview insights is better than a statistically perfect model that doesn’t make conceptual sense.
- Limit changes to 2-3 paths/items at a time—re-test fit after each adjustment to understand what’s driving improvements.
- Cross-validate: If your sample size allows, split it into training and testing sets—refine the model on the training set, then test fit on the testing set to ensure generalizability.
内容的提问来源于stack exchange,提问作者Janina Steinert

