实现筛选geom_bar结果并保留有效比例的分组分级条形图
解决方案
步骤说明
针对你的两个需求,我们先对数据进行预处理,再调整绘图逻辑:
- 按组计算患者百分比并拆分等级:以唯一
id-group为患者单位,统计每组内各问题不同等级的患者占比(占该组总患者数的百分比)。 - 筛选问题:先计算组1中每个问题的总患者占比,保留占比超过阈值
X的问题,同时保留这些问题在组2中的对应数据。
完整代码
library(tidyverse) library(scales) # 1. 设置阈值X(可根据需求调整) X <- 0.3 # 2. 数据预处理 data <- data.frame(id=c(rep(1,4),rep(2,8),rep(3,6)), group=c(1,1,1,2,1,1,1,2,2,2,2,2,1,1,2,2,2,2), problem=c(1,2,3,1,2,4,5,1,2,3,6,4,5,1,2,3,4,5), grade=c(rep(1,9),rep(2,9))) # 计算每组的总患者数(唯一id-group组合) group_total <- data |> distinct(id, group) |> count(group, name = "total_patients") # 统计每组-问题-等级的患者数及占比 problem_stats <- data |> distinct(id, group, problem, grade) |> # 确保每个患者的问题-等级唯一 count(group, problem, grade, name = "patient_count") |> left_join(group_total, by = "group") |> mutate(percentage = patient_count / total_patients) # 筛选组1中总占比超过X的问题 qualified_problems <- problem_stats |> filter(group == 1) |> group_by(problem) |> summarise(total_pct = sum(percentage)) |> filter(total_pct > X) |> pull(problem) # 保留符合条件的问题数据,并加入组1的总占比用于排序 filtered_data <- problem_stats |> filter(problem %in% qualified_problems) |> left_join( problem_stats |> filter(group == 1) |> group_by(problem) |> summarise(group1_total_pct = sum(percentage)), by = "problem" ) # 3. 绘图 ggplot(filtered_data, aes(x = fct_reorder(factor(problem), -group1_total_pct), y = percentage, fill = factor(grade), color = factor(group))) + # 绘制分组条形,内部按grade拆分 geom_col(aes(group = factor(group)), position = position_dodge(width = 0.8), width = 0.7, size = 1) + # 显示百分比标签 geom_text(aes(label = sprintf("%.0f%%", percentage * 100), group = factor(group)), position = position_dodge(width = 0.8), vjust = -0.3, size = 3.5) + # 坐标轴与配色设置(保持原代码风格) scale_y_continuous(labels = percent_format(accuracy = 1), expand = c(0, 0), limits = c(0, NA)) + scale_fill_grey(start = 0.6, end = 0.9) + scale_color_grey(start = 0.2, end = 0.5) + labs(x = NULL, y = "Percentage of Patients", fill = "Grade", color = "Group") + # 主题调整 theme_classic() + theme(axis.text.x = element_text(angle = 90, vjust = 0.5, hjust = 1), axis.ticks.x = element_blank(), legend.position = "bottom")
代码说明
- 数据处理部分:通过
distinct(id, group)确定每组总患者数,再统计每个问题-等级的患者占比,确保百分比计算的是患者占比而非记录占比。 - 筛选逻辑:仅保留组1中总占比超过
X的问题,同时保留这些问题在组2中的数据,满足你只展示关键问题的需求。 - 绘图部分:用
fct_reorder按组1的问题占比降序排列x轴,position_dodge实现同一问题的两组条形并排,条形内部按grade填充颜色,同时保留原代码的灰色配色风格。
内容的提问来源于stack exchange,提问作者scott9
相关产品推荐
相关产品推荐

