You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法绘制二分类因变量与连续自变量关联图?求解决方案

年龄与二分类结局关联的绘图问题

我想要展示年龄(V1,连续变量)与二分类结局(V2)的关联,但绘图时遇到问题:使用以下代码后未出现平滑曲线,期望生成类似示例图的效果,也欢迎更好的绘图方法建议。

数据集

> dput(head(test, 100))
structure(list(V1 = c(48, 92, 36, NA, 69, NA, NA, 19, 69, 82, 
NA, 39, 42, NA, 68, 72, 27, 78, 42, 15, 79, 48, 38, 46, 17, 33, 
24, 41, 68, 28, 79, NA, 52, 81, 74, 58, 57, 71, 51, 51, 51, 51, 
31, 96, 47, NA, 66, 66, 73, 55, 79, 60, 60, 76, 34, 53, 58, 70, 
80, 33, 17, 54, 42, 64, NA, 72, 53, 55, 59, NA, 68, 71, 70, 77, 
16, 74, 74, 29, 49, NA, 64, 65, 65, 65, 57, 63, 60, 78, 77, 75, 
54, 55, 97, NA, NA, 74, 80, 73, 74, 67), V2 = c(1, 0, 1, NA, 
1, NA, NA, 1, 1, 1, NA, 0, 1, NA, 1, 1, 1, 1, 1, 1, 1, 1, 0, 
1, 1, 1, 1, 0, 1, 1, 0, NA, 1, 0, 1, 1, 0, 0, 1, 1, 1, 1, 1, 
1, 1, NA, 1, 1, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 0, 1, 1, 
1, NA, 1, 1, 1, 1, NA, 0, 1, 1, 1, 1, 1, 0, 1, 0, NA, 1, 1, 1, 
1, 0, 0, 0, 1, 0, 1, 1, 0, 0, NA, NA, 0, 1, 0, 0, 0)), row.names = c(NA, 
100L), class = "data.frame")

尝试过的代码

ggplot(test, aes(x=V1, y=V2))+
  geom_point(size=2, alpha=0.4)+
  stat_smooth(method="loess", color="blue", size=1.5)

问题原因

V2是二分类变量(0/1),而loess是针对连续因变量的平滑拟合方法,直接使用会导致拟合逻辑错误,因此无法生成曲线;同时数据中存在的NA值也会干扰拟合过程。

解决方案

方法1:逻辑回归平滑(基础方案)

先过滤缺失值,使用逻辑回归拟合年龄与结局概率的关联曲线:

library(ggplot2)
# 过滤含NA的行
test_clean <- na.omit(test)

ggplot(test_clean, aes(x = V1, y = V2)) +
  geom_point(size = 2, alpha = 0.4) +
  # 用二项分布族的glm拟合概率曲线
  stat_smooth(method = "glm", 
              method.args = list(family = binomial),
              color = "blue", size = 1.5) +
  labs(x = "年龄", y = "结局概率(V2=1)")

方法2:广义可加模型(GAM,灵活非线性拟合)

如果需要更贴合数据的非线性曲线,可使用mgcv包的广义可加模型:

library(ggplot2)
library(mgcv)
test_clean <- na.omit(test)

ggplot(test_clean, aes(x = V1, y = V2)) +
  geom_point(size = 2, alpha = 0.4) +
  stat_smooth(method = "gam", 
              formula = y ~ s(x),  # s(x)表示对x做平滑处理
              method.args = list(family = binomial),
              color = "red", size = 1.5) +
  labs(x = "年龄", y = "结局概率(V2=1)")

方法3:分组比例拟合(直观展示)

将年龄分组后计算每组的结局比例,再拟合平滑曲线,更易解读:

library(ggplot2)
library(dplyr)
test_clean <- na.omit(test)

# 按年龄分组,计算每组的结局比例和年龄中点
test_grouped <- test_clean %>%
  mutate(age_group = cut(V1, breaks = seq(10, 100, by = 5))) %>%
  group_by(age_group) %>%
  summarise(
    outcome_ratio = mean(V2),
    age_midpoint = mean(V1, na.rm = TRUE)
  )

ggplot(test_grouped, aes(x = age_midpoint, y = outcome_ratio)) +
  # 保留原始散点
  geom_point(data = test_clean, aes(x = V1, y = V2), size = 2, alpha = 0.4) +
  # 添加分组比例点
  geom_point(size = 3, color = "orange") +
  # 对分组比例拟合平滑曲线
  stat_smooth(method = "loess", color = "blue", size = 1.5) +
  labs(x = "年龄", y = "结局比例(V2=1)")

内容的提问来源于stack exchange,提问作者Jamie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 20:40:30