R中BIC值不一致问题咨询:lme4线性混合效应模型对比
Hey there! Let's walk through what's happening with your linear mixed models and those confusing BIC values—this is a super common pitfall when working with lme4, so you're not alone.
1. Why do anova() and BIC() give different BIC values?
The key here is how lme4 handles model fitting for comparisons:
- When you fit a model with
lmer()without specifyingREML=FALSE, it uses REML (Restricted Maximum Likelihood) by default. TheBIC()function applied directly to this model will calculate the BIC based on the REML likelihood. - But when you run
anova(model1, model2)to compare models with different fixed effects,lme4automatically refits both models using ML (Maximum Likelihood) under the hood. This is because REML isn't valid for comparing models with different fixed effect structures—it's only meant for comparing changes to random effects.
So the discrepancy comes down to one set of BIC values being REML-based, the other ML-based. They're measuring slightly different things, hence the mismatch.
2. Why is BIC increasing when I add more factors?
First, let's clear up a common misconception: BIC is supposed to penalize model complexity. A higher BIC doesn't always mean you did something wrong—it might just mean your new factor isn't adding enough explanatory power to justify the extra parameter.
But if this feels unexpected, here are two likely reasons:
- You're comparing REML-based BIC values across models with different fixed effects: As mentioned earlier, this is invalid. REML's BIC doesn't account for changes in fixed effects properly, so these comparisons don't reflect true model fit vs. complexity tradeoffs.
- Even with ML-based BIC, the penalty for adding a new parameter outweighs the fit improvement: BIC balances how well the model fits the data against how many parameters it uses. If your new factor doesn't explain enough variance in the response, the penalty term (which scales with the number of parameters) will push BIC higher, telling you the simpler model is better.
3. How to correctly compare models with BIC
Follow these steps to get consistent, valid results:
- Fit all models with ML: Specify
REML=FALSEwhen building your models so you're working with comparable likelihoods:# Simple model model_simple <- lmer(y ~ fixed1 + (1|random_group), data = my_data, REML = FALSE) # Model with extra factor model_complex <- lmer(y ~ fixed1 + fixed2 + (1|random_group), data = my_data, REML = FALSE) - Compare with
anova(): Now the BIC values in theanova()output will match what you get fromBIC(model_simple)andBIC(model_complex):anova(model_simple, model_complex) # Check individual BIC values BIC(model_simple) BIC(model_complex) - Interpret the results: If
model_complexhas a lower BIC thanmodel_simple, the extra factor is worth keeping. If BIC is higher, stick with the simpler model—your data doesn't support the added complexity.
Quick Reminders
- Never use REML to compare models with different fixed effect structures. This is a hard rule in mixed model statistics.
- BIC's job is to prevent overfitting. A rising BIC when adding factors isn't a bug—it's the penalty doing its job!
内容的提问来源于stack exchange,提问作者Agata

