ggplot2绘制Logistic回归图呈线性而非S型曲线,求解决方法
问题分析与解决思路
核心原因
绘图模型与拟合模型不匹配:你拟合的是带个体随机效应
(1|p_id)的混合效应Logistic回归模型,但geom_smooth(method="glm")调用的是普通固定效应GLM,两者的参数估计和预测结果存在差异,并非你实际拟合的模型输出。S型曲线的视觉线性化:即使是正确的Logistic模型,当预测变量
imd_score的取值范围恰好落在S型曲线的中间近似线性段时,曲线会呈现出接近直线的形态。这种情况通常发生在:imd_score的回归系数绝对值很小,导致在0-80的范围内,预测概率的变化幅度始终处于S曲线的中间平缓上升/下降段;- 数据中
imd_score的取值集中在S型曲线的线性区间,未覆盖到概率趋近于0或1的极端区域。
解决步骤
步骤1:验证模型系数与数据分布
先查看你拟合的混合效应模型中imd_score的系数,确认其对概率的影响幅度:
summary(trees_imd_urb)
同时检查imd_score的分布和预测概率的范围:
# 查看IMD得分分布 hist(um_sub$imd_score) # 查看模型对原始数据的预测概率范围 um_sub$pred <- predict(trees_imd_urb, type = "response") range(um_sub$pred)
如果预测概率的范围在0.3-0.7之间,视觉上就会呈现线性。
步骤2:绘制实际混合效应模型的预测曲线
要得到与你拟合的glmer模型一致的曲线,需要手动生成预测数据并绘制:
- 创建包含
imd_score全范围、其他协变量取典型值的模拟数据集:
new_data <- expand.grid( imd_score = seq(0, 80, by = 1), # 连续变量取均值,分类变量取参考水平 age = mean(um_sub$age, na.rm = TRUE), gender = levels(um_sub$gender)[1], ethnicity = levels(um_sub$ethnicity)[1], education = levels(um_sub$education)[1], occupation = levels(um_sub$occupation)[1], urban_grow = levels(um_sub$urban_grow)[1], urban_rural = levels(um_sub$urban_rural)[1], p_id = unique(um_sub$p_id)[1] # 取单个个体或后续取平均 )
- 从
glmer模型中预测概率(re.form=NA表示只考虑固定效应的群体水平预测):
new_data$pred_prob <- predict(trees_imd_urb, newdata = new_data, type = "response", re.form = NA)
- 结合原始数据和预测结果绘图:
log_plot_trees_imd <- ggplot() + geom_jitter(data = um_sub, aes(x = imd_score, y = as.numeric(trees)-1), width = 0, height = 0.05) + geom_line(data = new_data, aes(x = imd_score, y = pred_prob), color = "darkblue", linewidth = 1) + labs(x = "IMD scores", y = "Probability of Seeing Trees") print(log_plot_trees_imd)
步骤3:(可选)展示随机效应的影响
如果需要展示个体随机效应的差异,可以保留随机效应预测,然后对多个个体的曲线取平均或绘制代表性个体的曲线:
# 生成包含多个p_id的预测数据 new_data_multi <- expand.grid( imd_score = seq(0, 80, by = 1), age = mean(um_sub$age, na.rm = TRUE), gender = levels(um_sub$gender)[1], ethnicity = levels(um_sub$ethnicity)[1], education = levels(um_sub$education)[1], occupation = levels(um_sub$occupation)[1], urban_grow = levels(um_sub$urban_grow)[1], urban_rural = levels(um_sub$urban_rural)[1], p_id = sample(unique(um_sub$p_id), 10) # 随机选10个个体 ) new_data_multi$pred_prob <- predict(trees_imd_urb, newdata = new_data_multi, type = "response") # 绘制个体曲线+群体平均曲线 log_plot_trees_imd <- ggplot() + geom_jitter(data = um_sub, aes(x = imd_score, y = as.numeric(trees)-1), width = 0, height = 0.05, alpha = 0.3) + geom_line(data = new_data_multi, aes(x = imd_score, y = pred_prob, group = p_id), color = "lightblue", linewidth = 0.8) + geom_line(data = new_data, aes(x = imd_score, y = pred_prob), color = "darkblue", linewidth = 1.2) + labs(x = "IMD scores", y = "Probability of Seeing Trees") print(log_plot_trees_imd)
内容的提问来源于stack exchange,提问作者Luca
相关产品推荐
相关产品推荐

