如何在ggplot2分面堆叠条形图中为同一Factor设置独立排序?
分面堆叠条形图实现分组独立排序(按"very low"占比降序)
问题背景
现有tibble数据集research_culture_by_career,包含调研响应ID、职业阶段(Career_stage)、工作相关因子(Factor)、感受等级(Level)。需要绘制按Career_stage分面的100%水平堆叠条形图,要求每个分面内的Factor按该分组下"very low"等级的占比独立降序排列,但默认分面设置无法实现该需求。
解决方案
方法1:基础实现(无需额外包)
通过给每个职业阶段下的Factor创建带分组前缀的唯一因子,配合scales = "free_y"实现独立排序:
library(tidyverse) # 读取数据 research_culture_by_career <- read_csv(file = "https://0x0.st/HrTp.csv", col_types = c("i", "c", "f", "f")) # 计算每个职业阶段下各Factor的very low占比,生成排序后的带分组因子 sorted_factors <- research_culture_by_career %>% group_by(Career_stage, Factor) %>% summarise( total = n(), very_low_count = sum(Level == "very low"), very_low_prop = very_low_count / total, .groups = "drop" ) %>% group_by(Career_stage) %>% arrange(desc(very_low_prop), .by_group = TRUE) %>% mutate(factor_with_group = paste(Career_stage, Factor, sep = "|")) %>% mutate(factor_with_group = fct_inorder(factor_with_group)) # 合并回原始数据集 research_culture_by_career <- research_culture_by_career %>% left_join(sorted_factors %>% select(Career_stage, Factor, factor_with_group), by = c("Career_stage", "Factor")) # 绘制分面图 research_culture_by_career_fig <- research_culture_by_career %>% group_by(Career_stage, factor_with_group, Level) %>% summarise(prop = n(), .groups = "drop") %>% ggplot(aes(x = prop, y = factor_with_group, fill = Level)) + geom_col(position = "fill") + facet_wrap(~Career_stage, scales = "free_y") + # 开启y轴独立刻度 labs(x = "Proportion of responses", y = NULL) + theme_classic() + scale_x_continuous(breaks = c(0, 0.5, 1.0), labels = scales::percent) + scale_fill_brewer(palette = "OrRd") + scale_y_discrete(labels = function(x) str_remove(x, "^.*\\|")) + # 移除分组前缀,显示原始Factor名称 theme(legend.position = "bottom") + guides(fill = guide_legend(title = NULL, nrow = 1, byrow = TRUE)) research_culture_by_career_fig
方法2:使用ggh4x包简化实现
ggh4x的facet_wrap2支持independent = TRUE参数,直接实现分面内因子独立排序,无需手动处理因子:
library(tidyverse) # install.packages("ggh4x") # 首次使用需安装 library(ggh4x) # 读取数据 research_culture_by_career <- read_csv(file = "https://0x0.st/HrTp.csv", col_types = c("i", "c", "f", "f")) # 按职业阶段分组,给Factor按very low占比降序排序 research_culture_by_career <- research_culture_by_career %>% group_by(Career_stage, Factor) %>% summarise(very_low_prop = sum(Level == "very low")/n(), .groups = "drop") %>% group_by(Career_stage) %>% arrange(desc(very_low_prop), .by_group = TRUE) %>% mutate(Factor = fct_inorder(Factor)) %>% right_join(research_culture_by_career, by = c("Career_stage", "Factor")) # 绘制分面图 research_culture_by_career_fig <- research_culture_by_career %>% group_by(Career_stage, Factor, Level) %>% summarise(prop = n(), .groups = "drop") %>% ggplot(aes(x = prop, y = Factor, fill = Level)) + geom_col(position = "fill") + facet_wrap2(~Career_stage, scales = "free_y", independent = TRUE) + # 开启独立排序 labs(x = "Proportion of responses", y = NULL) + theme_classic() + scale_x_continuous(breaks = c(0, 0.5, 1.0), labels = scales::percent) + scale_fill_brewer(palette = "OrRd") + theme(legend.position = "bottom") + guides(fill = guide_legend(title = NULL, nrow = 1, byrow = TRUE)) research_culture_by_career_fig
关键说明
- 方法1中
scales = "free_y"让每个分面的y轴独立显示,带分组前缀的因子确保每个分组内的排序逻辑不冲突; - 方法2中
facet_wrap2(independent = TRUE)直接支持分面内因子的独立排序,代码更简洁; - 两种方法都先计算了每个职业阶段下各Factor的"very low"占比,并以此为依据对Factor进行排序。
内容的提问来源于stack exchange,提问作者hpy
相关产品推荐
相关产品推荐

