如何让ggplot柱状图百分比反映组内占比而非整体占比
问题描述
我在R中有如下数据集:
(注:f是数据集,列e为字符型分组变量,列d为因子型数据列)
使用以下代码生成了图表:
ggplot(subset(f,!is.na(d)),aes(x=d,y=(..count..)/sum(..count..),fill=forcats::fct_rev(e))) + geom_bar(position="dodge") + scale_y_continuous(labels = scales::percent) + theme(panel.grid.major.y = element_line(color="gray"), panel.background =element_blank(), axis.line=element_line("black"), axis.text.x=element_text(face="bold"), legend.title=element_blank(), axis.title.y=element_blank(), axis.title.x=element_blank(), plot.title = element_text(hjust = 0.5) ) + scale_x_discrete( labels=c('$0', '$1 to $10K', 'Over $10K to $20K', 'Over $20K to $30K','Over $30K to $40K','Over $40K to $50K','Over $50K') ) + geom_text(aes(label = scales::percent(round((..count..)/sum(..count..),2)),y= ((..count..)/sum(..count..))), stat="count", position = position_dodge(width = .9),vjust = -1) labs(title="Perceived Debt Before and After CSP 2024") + scale_fill_manual(values=c("darkblue","lightblue"))
生成的图表如下:
目前图表中两组的百分比总和为100%,但我希望每组(Pre-CSP和Post-CSP)内的百分比总和为100%,请问该如何实现?
注:X轴已重新标注,数据列d中的1代表$0,2代表'$1 to $10,000',以此类推。
解决方案
要实现每组(e列的Pre-CSP/Post-CSP)内百分比总和为100%,核心是按分组变量e计算每组内的占比,而非全局占比。以下是两种可行方案:
方案一:提前用dplyr预处理计算占比(推荐,逻辑清晰)
先对数据做分组统计,计算每组内的百分比,再用ggplot绘图:
library(dplyr) library(ggplot2) # 预处理数据:按分组和类别统计数量,再计算组内占比 f_processed <- subset(f, !is.na(d)) %>% group_by(e, d) %>% summarise(count = n(), .groups = "drop") %>% group_by(e) %>% mutate(pct = count / sum(count)) %>% ungroup() %>% mutate(e = forcats::fct_rev(e)) # 保持原代码的因子反转逻辑 # 绘图 ggplot(f_processed, aes(x = d, y = pct, fill = e)) + geom_bar(stat = "identity", position = "dodge") + scale_y_continuous(labels = scales::percent) + theme( panel.grid.major.y = element_line(color = "gray"), panel.background = element_blank(), axis.line = element_line("black"), axis.text.x = element_text(face = "bold"), legend.title = element_blank(), axis.title.y = element_blank(), axis.title.x = element_blank(), plot.title = element_text(hjust = 0.5) ) + scale_x_discrete( labels = c('$0', '$1 to $10K', 'Over $10K to $20K', 'Over $20K to $30K', 'Over $30K to $40K','Over $40K to $50K','Over $50K') ) + geom_text(aes(label = scales::percent(round(pct, 2))), position = position_dodge(width = .9), vjust = -1) + labs(title = "Perceived Debt Before and After CSP 2024") + scale_fill_manual(values = c("darkblue", "lightblue"))
方案二:在ggplot内直接按分组计算占比(无需预处理)
修改原代码的aes映射,用ave()函数按e分组计算组内总数,从而得到组内占比:
ggplot(subset(f, !is.na(d)), aes(x = d, fill = forcats::fct_rev(e))) + geom_bar(aes(y = stat(count) / ave(count, group = e)), position = "dodge") + scale_y_continuous(labels = scales::percent) + theme( panel.grid.major.y = element_line(color = "gray"), panel.background = element_blank(), axis.line = element_line("black"), axis.text.x = element_text(face = "bold"), legend.title = element_blank(), axis.title.y = element_blank(), axis.title.x = element_blank(), plot.title = element_text(hjust = 0.5) ) + scale_x_discrete( labels = c('$0', '$1 to $10K', 'Over $10K to $20K', 'Over $20K to $30K', 'Over $30K to $40K','Over $40K to $50K','Over $50K') ) + geom_text(aes(y = stat(count) / ave(count, group = e), label = scales::percent(round(stat(count)/ave(count, group = e), 2))), stat = "count", position = position_dodge(width = .9), vjust = -1) + labs(title = "Perceived Debt Before and After CSP 2024") + scale_fill_manual(values = c("darkblue", "lightblue"))
关键改动说明
- 方案一通过dplyr提前计算占比,便于数据检查和后续调整,逻辑更直观;
- 方案二在ggplot内部利用
ave(count, group = e)实现分组求和,直接得到组内占比,无需额外数据处理; - 两种方案都能让Pre-CSP和Post-CSP各自组内的百分比总和达到100%。
内容的提问来源于stack exchange,提问作者Mike Wang
相关产品推荐
相关产品推荐

