You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言ggplot分组柱状图:同百分比为何柱高不同?

分组柱状图百分比展示问题与解决方法

问题描述

需要在study分组内展示more_know(1-7类别)的百分比,绘制分组柱状图时,两类study的more_know=1类别文本均显示3%,但柱子高度存在差异,推测是精确值与四舍五入后的数值差异导致,需调整绘制方案实现合理的可视化效果。

原始数据

labels.feed2 <- c(1:7)

df.sci.cin <- data.frame(
  study = factor(c(rep(1,62), rep(2,33)), levels=1:2),
  more_know = factor(c(6, 2, 2, 2, 4, 5, 5, 3, 4, 5, 5, 5, 4, 2, 4, 7, 7, 2, 7, 5, 5, 5, 6, 2, 4, 7, 2, 5, 3, 2, 5, 7, 3, 5, 4, 4, 5, 4, 6, 5, 5, 7, 5, 1, 5, 5, 2, 4, 2, 7, 5, 5, 2, 5, 4, 6, 5, 7, 1, 5, 4, 3, 5, 4, 5, 2, 5, 6, 5, 3, 2, 2, 6, 2, 4, 5, 2, 5, 3, 5, 7, 7, 4, 5, 6, 3, 3, 1, 5, 4, 4, 6, 6, 4, 4), levels=1:7, labels=labels.feed2)
)

tibble.sci.cin <- as.tibble(df.sci.cin)

原始绘图代码

vec.labels.more_know <- c("1 
familiar with 
all of this", 2:6, "7 
very new 
to me")

ggplot(data = tibble.sci.cin, aes(   
  x = factor(more_know, levels = 1:7, labels = vec.labels.more_know),
  fill = factor(study)
)) +
  geom_bar(
    aes(
      y = after_stat(count / ave(count, fill, FUN = sum))
    ),
    position = "dodge"
  ) +
  scale_fill_manual(
    values = c("grey40", "grey60"),
    name = "event location",
    labels = c("university (n=62)", "cinema (n=33)")
  ) +
  geom_text(
    aes(
      y = after_stat(count / ave(count, fill, FUN = sum)),
      label = after_stat(scales::percent(count / ave(count, fill, FUN = sum), accuracy = 1))
    ),
    stat = "count", position = position_dodge(0.9), vjust = -0.5
  ) +
  ylab("percent of audience relative to location") +
  xlab("feeling of more knowledge of climate change after the event") +
  theme(axis.text.x = element_text(hjust = .9)) + #angle = 45, 
  theme(axis.ticks.x = element_blank()) +
  scale_y_continuous(labels = scales::percent, limits = c(0, 0.38)) +
  scale_x_discrete(drop = FALSE) +
  theme(
    panel.border = element_rect(linetype = "solid", colour = "black", linewidth = .5, fill = NA),
    panel.grid.minor = element_line(colour = "grey93", linewidth = .3),
    panel.grid.major.y = element_line(colour = "grey93", linewidth = .3),
    panel.background = element_rect(fill = "grey97")
    ) +
  theme(axis.title.x.bottom = element_text(margin = margin(t = .15, unit = "in")))

问题原因

计算两类study中more_know=1的精确百分比:

  • study=1(样本量62):2/62≈3.226%
  • study=2(样本量33):1/33≈3.030%
    两者四舍五入后均显示为3%,但实际精确值存在差异,因此柱子高度不同。

解决方案

方案1:提前预处理数据(推荐,兼顾准确性与可视化)

不在ggplot中实时计算百分比,先通过数据预处理得到每个分组的精确百分比,统一控制文本显示与柱子高度:

library(dplyr)
library(ggplot2)

# 预处理计算百分比与标签
tibble.sci.cin_summary <- tibble.sci.cin %>%
  count(study, more_know) %>%
  group_by(study) %>%
  mutate(
    percent = n / sum(n),
    percent_label = scales::percent(percent, accuracy = 1)
  ) %>%
  ungroup()

# 定义x轴标签
vec.labels.more_know <- c("1 
familiar with 
all of this", 2:6, "7 
very new 
to me")

优化后的绘图代码:

ggplot(data = tibble.sci.cin_summary, aes(
  x = factor(more_know, levels = 1:7, labels = vec.labels.more_know),
  y = percent,
  fill = factor(study)
)) +
  geom_col(position = position_dodge(0.9)) +
  scale_fill_manual(
    values = c("grey40", "grey60"),
    name = "event location",
    labels = c("university (n=62)", "cinema (n=33)")
  ) +
  geom_text(
    aes(label = percent_label),
    position = position_dodge(0.9),
    vjust = -0.5,
    size = 3.5
  ) +
  ylab("percent of audience relative to location") +
  xlab("feeling of more knowledge of climate change after the event") +
  theme(axis.text.x = element_text(hjust = .9)) +
  theme(axis.ticks.x = element_blank()) +
  scale_y_continuous(labels = scales::percent, limits = c(0, 0.38)) +
  scale_x_discrete(drop = FALSE) +
  theme(
    panel.border = element_rect(linetype = "solid", colour = "black", linewidth = .5, fill = NA),
    panel.grid.minor = element_line(colour = "grey93", linewidth = .3),
    panel.grid.major.y = element_line(colour = "grey93", linewidth = .3),
    panel.background = element_rect(fill = "grey97")
  ) +
  theme(axis.title.x.bottom = element_text(margin = margin(t = .15, unit = "in")))

方案2:强制柱子高度一致(仅用于视觉需求)

若需让more_know=1的柱子高度完全一致且文本显示3%,可在预处理时强制该类别的百分比为0.03,但此方法会改变数据精确性,不适合严谨统计场景:

tibble.sci.cin_summary <- tibble.sci.cin %>%
  count(study, more_know) %>%
  group_by(study) %>%
  mutate(
    percent = ifelse(more_know == 1, 0.03, n / sum(n)),
    percent_label = scales::percent(percent, accuracy = 1)
  ) %>%
  ungroup()

内容的提问来源于stack exchange,提问作者Hildrun Walter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 10:05:59