添加协变量后lmer模型ggplot可视化折线错乱的解决方法
问题原因
添加协变量后折线错乱的核心原因是直接用原始数据集dat生成预测值:
- 原始数据中
sex_dv、ethnicity、parentEducation等协变量存在多种取值,即使time_years和educ相同,不同协变量组合会输出不同的预测值。 geom_line会按time_years的顺序连接所有观测点,同一x值对应多个y值时,线条就会出现无意义的交叉、抖动。- 未调整模型没有这些协变量,同一
time_years+educ组合的预测值唯一,因此线条平滑。
另外原代码存在一处错误:计算置信区间时调用了dat$se_predicted,但该变量未被赋值,实际应使用predictions$se.fit。
解决方案
正确做法是创建参考预测数据集,固定协变量到参考水平(分类变量取众数、连续变量取均值),仅让time_years和educ按需求范围变化,确保每个组合仅生成一个预测值,线条自然平滑。
步骤1:创建参考数据集
# 提取协变量的参考水平 ref_sex <- names(table(dat$sex_dv))[which.max(table(dat$sex_dv))] ref_ethnicity <- names(table(dat$ethnicity))[which.max(table(dat$ethnicity))] ref_parentEdu <- names(table(dat$parentEducation))[which.max(table(dat$parentEducation))] # 生成覆盖time_years全范围的序列,结合educ的所有取值 time_seq <- seq(min(dat$time_years), max(dat$time_years), length.out = 100) pred_dat <- expand.grid( time_years = time_seq, educ = unique(dat$educ), before_after = 0, # 根据研究假设固定到参考水平 time_since_job = 0, # 同理固定到参考水平 sex_dv = ref_sex, ethnicity = ref_ethnicity, parentEducation = ref_parentEdu, pidp = NA # re.form=NA时无需分组变量 )
步骤2:生成预测值与置信区间
# 从调整后模型生成预测 predictions <- predict(lmer4, newdata = pred_dat, re.form = NA, type = "response", se.fit = TRUE) # 赋值预测值与95%CI pred_dat$predicted_values <- predictions$fit pred_dat$lower_bound <- pred_dat$predicted_values - 1.96 * predictions$se.fit pred_dat$upper_bound <- pred_dat$predicted_values + 1.96 * predictions$se.fit
步骤3:使用参考数据集绘图
ggplot(pred_dat, aes(x = time_years, y = predicted_values, color = as.factor(educ))) + theme_classic() + geom_line(linewidth = 1) + # 按教育分组绘制置信区间,避免区间重叠混淆 geom_ribbon(aes(ymin = lower_bound, ymax = upper_bound, fill = as.factor(educ)), alpha = 0.3, linetype = 0) + geom_vline(xintercept = 0, linetype = "dashed", color = "red") + scale_color_manual(values = c("blue", "red"), labels = c("No Degree", "Degree or higher")) + scale_fill_manual(values = c("blue", "red"), labels = c("No Degree", "Degree or higher")) + labs(x = "Time relative to job start (years)", y = "Daily Exercise", color = "Education Status", fill = "Education Status")
关键说明
- 固定协变量时需结合研究目的:若展示整体平均水平,分类变量取众数、连续变量取均值;若展示特定亚组,直接固定协变量到对应取值即可。
before_after和time_since_job也需固定到参考水平,否则其取值波动会导致预测值异常。
内容的提问来源于stack exchange,提问作者Rnewbie
相关产品推荐
相关产品推荐

