You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

添加协变量后lmer模型ggplot可视化折线错乱的解决方法

问题原因

添加协变量后折线错乱的核心原因是直接用原始数据集dat生成预测值:

  • 原始数据中sex_dv、ethnicity、parentEducation等协变量存在多种取值,即使time_years和educ相同,不同协变量组合会输出不同的预测值。
  • geom_line会按time_years的顺序连接所有观测点,同一x值对应多个y值时,线条就会出现无意义的交叉、抖动。
  • 未调整模型没有这些协变量,同一time_years+educ组合的预测值唯一,因此线条平滑。

另外原代码存在一处错误:计算置信区间时调用了dat$se_predicted,但该变量未被赋值,实际应使用predictions$se.fit。

解决方案

正确做法是创建参考预测数据集,固定协变量到参考水平(分类变量取众数、连续变量取均值),仅让time_years和educ按需求范围变化,确保每个组合仅生成一个预测值,线条自然平滑。

步骤1:创建参考数据集

# 提取协变量的参考水平
ref_sex <- names(table(dat$sex_dv))[which.max(table(dat$sex_dv))]
ref_ethnicity <- names(table(dat$ethnicity))[which.max(table(dat$ethnicity))]
ref_parentEdu <- names(table(dat$parentEducation))[which.max(table(dat$parentEducation))]

# 生成覆盖time_years全范围的序列,结合educ的所有取值
time_seq <- seq(min(dat$time_years), max(dat$time_years), length.out = 100)
pred_dat <- expand.grid(
  time_years = time_seq,
  educ = unique(dat$educ),
  before_after = 0,  # 根据研究假设固定到参考水平
  time_since_job = 0,  # 同理固定到参考水平
  sex_dv = ref_sex,
  ethnicity = ref_ethnicity,
  parentEducation = ref_parentEdu,
  pidp = NA  # re.form=NA时无需分组变量
)

步骤2:生成预测值与置信区间

# 从调整后模型生成预测
predictions <- predict(lmer4, newdata = pred_dat, re.form = NA, type = "response", se.fit = TRUE)

# 赋值预测值与95%CI
pred_dat$predicted_values <- predictions$fit
pred_dat$lower_bound <- pred_dat$predicted_values - 1.96 * predictions$se.fit
pred_dat$upper_bound <- pred_dat$predicted_values + 1.96 * predictions$se.fit

步骤3:使用参考数据集绘图

ggplot(pred_dat, aes(x = time_years, y = predicted_values, color = as.factor(educ))) +
  theme_classic() +
  geom_line(linewidth = 1) +
  # 按教育分组绘制置信区间,避免区间重叠混淆
  geom_ribbon(aes(ymin = lower_bound, ymax = upper_bound, fill = as.factor(educ)), 
              alpha = 0.3, linetype = 0) +
  geom_vline(xintercept = 0, linetype = "dashed", color = "red") +
  scale_color_manual(values = c("blue", "red"), 
                     labels = c("No Degree", "Degree or higher")) +
  scale_fill_manual(values = c("blue", "red"), 
                    labels = c("No Degree", "Degree or higher")) +
  labs(x = "Time relative to job start (years)", 
       y = "Daily Exercise", 
       color = "Education Status",
       fill = "Education Status")
关键说明
  • 固定协变量时需结合研究目的:若展示整体平均水平,分类变量取众数、连续变量取均值;若展示特定亚组,直接固定协变量到对应取值即可。
  • before_after和time_since_job也需固定到参考水平,否则其取值波动会导致预测值异常。

内容的提问来源于stack exchange,提问作者Rnewbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 11:17:34