ggplot2 geom_line绘制随访用药数据:NA值未中断线条问题
问题:随访用药绘图时NA值未中断线条且出现异常连线
我有随访数据,想按受试者绘制用药情况,但使用geom_line绘图时,NA值位置线条未中断,且第3个分面出现了不该有的药物线条。期望效果:顶部线条表示随访天数,下方线条展示对应日期的用药情况。
数据示例
# A tibble: 23 × 3 ID day medication <dbl> <dbl> <chr> 1 1 1 NA 2 1 2 A 3 1 3 NA 4 1 4 A 5 1 5 A 6 1 5 B 7 1 6 A 8 2 1 NA 9 2 2 C 10 2 3 NA 11 3 1 A 12 3 2 C 13 3 2 D 14 3 3 D 15 3 4 NA 16 3 5 A 17 4 1 C 18 4 1 D 19 4 2 C 20 4 2 D 21 4 3 C 22 4 3 D 23 4 4 NA
当前绘图代码
k <- df %>% ggplot() + geom_line(aes(x=day, y= medication ,color= medication, group=medication),linewidth = 1,) + facet_grid(ID~ ., drop = TRUE, scales = "free", space = "free")
问题根源
- 将分类变量
medication直接作为y轴值,ggplot会自动将其转为连续整数,导致NA与非NA值之间被错误连线 group=medication仅按药物分组,会让同一药物跨不同受试者、跨NA值强制连接,引发异常线条
解决方案
通过数据整理+合理分组,实现NA值处线条中断,并添加随访天数顶线:
library(tidyverse) # 整理数据:补全每个受试者-药物的所有随访天数,无用药日期留NA;同时添加随访顶线数据 df_plot <- df %>% # 过滤NA并给药物分配y轴位置 filter(!is.na(medication)) %>% mutate(y = as.integer(factor(medication))) %>% # 补全每个ID-药物组合的所有天数,无用药则y为NA group_by(ID, medication) %>% complete(day = full_seq(day, 1)) %>% ungroup() %>% # 添加随访天数顶线数据(y值高于所有药物) bind_rows( df %>% group_by(ID) %>% summarise( day = c(1, max(day)), medication = "随访天数", y = max(as.integer(factor(medication)), na.rm = TRUE) + 1, .groups = "drop" ) ) # 绘图 ggplot(df_plot) + geom_line( aes(x = day, y = y, color = medication, group = interaction(ID, medication)), linewidth = 1, na.rm = FALSE ) + # 自定义颜色 scale_color_manual( values = c("A" = "#E64B35", "B" = "#4DBBD5", "C" = "#00A087", "D" = "#3C5488", "随访天数" = "#2C3E50") ) + # 分面设置,y轴自由缩放 facet_grid(ID~., scales = "free_y", space = "free_y") + labs(x = "随访天数", color = "类型") + theme_minimal() + theme(axis.title.y = element_blank())
效果说明
interaction(ID, medication)确保每个受试者的每个药物单独分组,不会跨ID或药物连线complete(day = full_seq(day, 1))补全所有天数,无用药日期的NA会让geom_line自动中断线条- 随访顶线单独生成,y值高于所有药物,清晰标记每个受试者的随访周期
内容的提问来源于stack exchange,提问作者ghs101
相关产品推荐
相关产品推荐

