You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

堆叠百分比柱状图各组样本量的直观可视化方案咨询

实现直观展示样本量的堆叠百分比柱状图(含子区段分割方案)

一、实现你设想的「子区段分割」方案

核心思路是将每组数据按样本数n拆分为对应数量的行,让每个子区段对应一个独立样本,再通过堆叠柱状图展示百分比占比,同时用小单元直观体现样本数量。

代码示例:

library(dplyr)
library(ggplot2)

# 加载数据并做基础汇总
df <- dplyr::starwars
df_sum <- na.omit(df %>% count(gender, sex))

# 扩展数据:把每组拆成n个独立行(每行对应一个样本)
df_expanded <- df_sum %>% uncount(n)

# 绘制带样本子区段的堆叠百分比柱状图
ggplot(df_expanded, aes(x = 1, y = gender, fill = sex)) +
  geom_bar(position = "fill", stat = "identity", width = 0.8, color = "white", size = 0.1) +
  scale_x_continuous(breaks = NULL) + # 隐藏x轴刻度,聚焦百分比占比
  labs(x = "百分比占比", y = "性别分类", fill = "生理性别") +
  theme_minimal()

效果说明:每个颜色段里的白色细格对应单个样本,既能看到各组的百分比占比,也能通过格子数量直观判断样本数多少(比如masculine, male的60个样本会显示为60个小单元)。


二、其他直观展示样本量的方案

1. 堆叠柱+样本数标注

在原百分比堆叠图上直接标注每个分段的样本数,同时补充组总样本数,兼顾百分比和绝对数量:

# 计算每组总样本数
df_total <- df_sum %>% 
  group_by(gender) %>% 
  summarize(total = sum(n)) %>% 
  ungroup()

df_sum <- df_sum %>% left_join(df_total, by = "gender")

ggplot(df_sum, aes(fill = sex, y = gender, x = n)) +
  geom_bar(position = "fill", stat = "identity", width = 0.8) +
  geom_text(aes(label = n), position = position_fill(vjust = 0.5), size = 3) +
  geom_text(data = df_total, aes(x = 1.1, label = paste("总样本:", total)), 
            inherit.aes = FALSE, size = 3) +
  scale_x_continuous(limits = c(0, 1.2)) + # 预留空间放总样本数
  labs(x = "百分比占比", y = "性别分类", fill = "生理性别") +
  theme_minimal()

2. 样本点分布图

用散点直接展示每个样本,按分组排列,直观性拉满:

ggplot(df_expanded, aes(x = sex, y = gender, color = sex)) +
  geom_jitter(width = 0.2, height = 0.1, size = 2, alpha = 0.8) +
  labs(x = "生理性别", y = "性别分类", color = "生理性别") +
  theme_minimal()

3. 双轴组合图(谨慎使用)

同时展示绝对样本数和百分比占比,但双轴容易造成视觉误导,需配合清晰标注:

max_total <- max(df_total$total)

ggplot(df_sum, aes(y = gender)) +
  # 左侧:绝对样本数堆叠柱
  geom_bar(aes(x = n, fill = sex), stat = "identity", width = 0.4, position = "stack") +
  # 右侧:百分比堆叠轮廓
  geom_bar(aes(x = n / max_total * 100), fill = NA, color = "black", 
           stat = "identity", width = 0.4, position = "fill") +
  scale_x_continuous(
    name = "样本数量",
    sec.axis = sec_axis(~ . / 100 * max_total, name = "百分比(%)")
  ) +
  labs(fill = "生理性别") +
  theme_minimal()

内容的提问来源于stack exchange,提问作者Patrick Rosendahl Andreassen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 07:05:11