如何用ggplot的geom_col按分组计算行百分比并绘制得分分布图
计算人口统计维度下得分等级的行百分比并绘制条形图
核心思路
按目标人口统计维度(种族/性别/教育水平)分组,统计每个维度组内各得分等级的样本数,再计算组内占比(行百分比),最后用ggplot2的geom_col绘制条形图。
步骤1:加载依赖包
library(dplyr) library(ggplot2)
步骤2:计算单维度行百分比
假设你的数据框名为df,包含race(种族)、gender(性别)、education(教育水平)、score(得分等级)列。以种族为例,计算行百分比:
race_percent <- df %>% # 按种族+得分等级分组 group_by(race, score) %>% # 统计每个组合的样本数,保留种族分组 summarize(count = n(), .groups = "drop_last") %>% # 计算每个得分在对应种族内的占比(转成百分比) mutate(percent = count / sum(count) * 100) %>% # 取消分组,避免后续操作干扰 ungroup()
同理,计算性别、教育水平维度的行百分比,只需替换group_by里的分组变量:
# 性别维度 gender_percent <- df %>% group_by(gender, score) %>% summarize(count = n(), .groups = "drop_last") %>% mutate(percent = count / sum(count) * 100) %>% ungroup() # 教育水平维度 education_percent <- df %>% group_by(education, score) %>% summarize(count = n(), .groups = "drop_last") %>% mutate(percent = count / sum(count) * 100) %>% ungroup()
步骤3:合并多维度数据(可选)
如果需要在同一张图里展示所有维度(比如分面),可以统一列名后合并数据:
combined_percent <- dplyr::bind_rows( race_percent %>% rename(group = race) %>% mutate(dimension = "种族"), gender_percent %>% rename(group = gender) %>% mutate(dimension = "性别"), education_percent %>% rename(group = education) %>% mutate(dimension = "教育水平") )
步骤4:绘制条形图
单维度绘图(以种族为例)
ggplot(race_percent, aes(x = factor(score), y = percent, fill = factor(race))) + geom_col(position = "dodge") + # 分组并列条形图 labs(x = "得分等级", y = "占比(%)", fill = "种族") + theme_minimal()
多维度分面绘图
ggplot(combined_percent, aes(x = factor(score), y = percent, fill = factor(group))) + geom_col(position = "dodge") + facet_wrap(~dimension, scales = "free") + # 按维度分面展示 labs(x = "得分等级", y = "占比(%)", fill = "分组") + theme_minimal()
关键说明
.groups = "drop_last":确保sum(count)计算的是每个人口统计分组的总样本数,而非全局总数factor(score):将得分等级转成分类变量,避免ggplot将其视为连续轴position = "dodge":让同得分等级下的不同分组条形并列展示,而非堆叠
内容的提问来源于stack exchange,提问作者a_todd12
相关产品推荐
相关产品推荐

