You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让ggplot柱状图百分比反映组内占比而非整体占比

问题描述

我在R中有如下数据集:
数据集截图
(注:f是数据集,列e为字符型分组变量,列d为因子型数据列)

使用以下代码生成了图表:

ggplot(subset(f,!is.na(d)),aes(x=d,y=(..count..)/sum(..count..),fill=forcats::fct_rev(e))) + 
 geom_bar(position="dodge") + 
 scale_y_continuous(labels = scales::percent) +
 theme(panel.grid.major.y = element_line(color="gray"),
    panel.background =element_blank(), axis.line=element_line("black"), axis.text.x=element_text(face="bold"), 
    legend.title=element_blank(), axis.title.y=element_blank(), axis.title.x=element_blank(),
    plot.title = element_text(hjust = 0.5)
 ) +
 scale_x_discrete(
   labels=c('$0', '$1 to $10K', 'Over $10K to $20K', 'Over $20K to $30K','Over $30K to $40K','Over $40K to $50K','Over $50K')
 ) +
 geom_text(aes(label = scales::percent(round((..count..)/sum(..count..),2)),y= ((..count..)/sum(..count..))), stat="count", position = position_dodge(width = .9),vjust = -1)
 labs(title="Perceived Debt Before and After CSP 2024") + 
 scale_fill_manual(values=c("darkblue","lightblue"))

生成的图表如下:
生成的图表

目前图表中两组的百分比总和为100%,但我希望每组(Pre-CSP和Post-CSP)内的百分比总和为100%,请问该如何实现?
注:X轴已重新标注,数据列d中的1代表$0,2代表'$1 to $10,000',以此类推。


解决方案

要实现每组(e列的Pre-CSP/Post-CSP)内百分比总和为100%,核心是按分组变量e计算每组内的占比,而非全局占比。以下是两种可行方案:

方案一:提前用dplyr预处理计算占比(推荐,逻辑清晰)

先对数据做分组统计,计算每组内的百分比,再用ggplot绘图:

library(dplyr)
library(ggplot2)

# 预处理数据:按分组和类别统计数量,再计算组内占比
f_processed <- subset(f, !is.na(d)) %>%
  group_by(e, d) %>%
  summarise(count = n(), .groups = "drop") %>%
  group_by(e) %>%
  mutate(pct = count / sum(count)) %>%
  ungroup() %>%
  mutate(e = forcats::fct_rev(e)) # 保持原代码的因子反转逻辑

# 绘图
ggplot(f_processed, aes(x = d, y = pct, fill = e)) +
  geom_bar(stat = "identity", position = "dodge") +
  scale_y_continuous(labels = scales::percent) +
  theme(
    panel.grid.major.y = element_line(color = "gray"),
    panel.background = element_blank(),
    axis.line = element_line("black"),
    axis.text.x = element_text(face = "bold"),
    legend.title = element_blank(),
    axis.title.y = element_blank(),
    axis.title.x = element_blank(),
    plot.title = element_text(hjust = 0.5)
  ) +
  scale_x_discrete(
    labels = c('$0', '$1 to $10K', 'Over $10K to $20K', 'Over $20K to $30K',
               'Over $30K to $40K','Over $40K to $50K','Over $50K')
  ) +
  geom_text(aes(label = scales::percent(round(pct, 2))), 
            position = position_dodge(width = .9), vjust = -1) +
  labs(title = "Perceived Debt Before and After CSP 2024") +
  scale_fill_manual(values = c("darkblue", "lightblue"))

方案二:在ggplot内直接按分组计算占比(无需预处理)

修改原代码的aes映射,用ave()函数按e分组计算组内总数,从而得到组内占比:

ggplot(subset(f, !is.na(d)), aes(x = d, fill = forcats::fct_rev(e))) +
  geom_bar(aes(y = stat(count) / ave(count, group = e)), position = "dodge") +
  scale_y_continuous(labels = scales::percent) +
  theme(
    panel.grid.major.y = element_line(color = "gray"),
    panel.background = element_blank(),
    axis.line = element_line("black"),
    axis.text.x = element_text(face = "bold"),
    legend.title = element_blank(),
    axis.title.y = element_blank(),
    axis.title.x = element_blank(),
    plot.title = element_text(hjust = 0.5)
  ) +
  scale_x_discrete(
    labels = c('$0', '$1 to $10K', 'Over $10K to $20K', 'Over $20K to $30K',
               'Over $30K to $40K','Over $40K to $50K','Over $50K')
  ) +
  geom_text(aes(y = stat(count) / ave(count, group = e), 
                label = scales::percent(round(stat(count)/ave(count, group = e), 2))),
            stat = "count", position = position_dodge(width = .9), vjust = -1) +
  labs(title = "Perceived Debt Before and After CSP 2024") +
  scale_fill_manual(values = c("darkblue", "lightblue"))

关键改动说明

  • 方案一通过dplyr提前计算占比,便于数据检查和后续调整,逻辑更直观;
  • 方案二在ggplot内部利用ave(count, group = e)实现分组求和,直接得到组内占比,无需额外数据处理;
  • 两种方案都能让Pre-CSP和Post-CSP各自组内的百分比总和达到100%。

内容的提问来源于stack exchange,提问作者Mike Wang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 03:43:18