如何用ggplot实现堆叠条形图的分组嵌套降序排序?
问题描述
我需要绘制6个年龄组按疾病类型划分的发病率堆叠条形图,要求每个年龄组内的疾病类型按value值降序排列(数值较大的类别位于条形底部)。我尝试了以下方法,但仅能按各年龄组的最大疾病类别排序,其余类别未正确排序,请问该需求是否可实现?
数据
structure(list(age = structure(c(7L, 7L, 8L, 8L, 9L, 9L, 10L, 10L, 11L, 11L, 12L, 12L, 7L, 7L, 8L, 8L, 9L, 9L, 10L, 10L, 11L, 11L, 12L, 12L, 7L, 7L, 8L, 8L, 9L, 9L, 10L, 10L, 11L, 11L, 12L, 12L, 7L, 7L, 8L, 8L, 9L, 9L, 10L, 10L, 11L, 11L, 12L, 12L, 7L, 7L, 8L, 8L, 9L, 9L, 10L, 10L, 11L, 11L, 12L, 12L, 7L, 7L, 8L, 8L, 9L, 9L, 10L, 10L, 11L, 11L, 12L, 12L, 7L, 7L, 8L, 8L, 9L, 9L, 10L, 10L, 11L, 11L, 12L, 12L, 7L, 7L, 8L, 8L, 9L, 9L, 10L, 10L, 11L, 11L, 12L, 12L, 7L, 7L, 8L, 8L, 9L, 9L, 10L, 10L, 11L, 11L, 12L, 12L, 7L, 7L, 8L, 8L, 9L, 9L, 10L, 10L, 11L, 11L, 12L, 12L), levels = c("<5 years", "5-9 years", "10-14 years", "15-19 years", "20-24 years", "25-29 years", "30-34 years", "35-39 years", "40-44 years", "45-49 years", "50-54 years", "55-59 years", "60-64 years", "65-69 years", "70-74 years", "75-79 years", "80-84 years", "85+ years", "Age-standardized"), class = "factor"), disease = c("infection", "infection", "infection", "infection", "infection", "infection", "infection", "infection", "infection", "infection", "infection", "infection", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "stroke", "infection", "infection", "infection", "infection", "infection", "infection", "infection", "infection", "infection", "infection", "infection", "infection", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "HBP", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "CKD", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea", "diarrhoea"), value = c(23353.7, 3.9, 23642.5, 4.4, 24538.1, 5, 28653.1, 6, 37310.1, 8.5, 46492.1, 12.5, 62458.4, 10.4, 90531.6, 16.7, 119950.1, 24.3, 154233.1, 32.6, 191329, 43.8, 225451.4, 60.8, 121.5, 5, 214.7, 8.8, 404.1, 17.2, 733.9, 31.2, 1397.4, 56.5, 2578.9, 95.2, 120830.8, 20.1, 188983.9, 34.9, 305603.5, 61.9, 480027.2, 101.3, 759052.9, 173.8, 1048404.1, 282.6, 1.7, 0.1, 2.2, 0.1, 3.3, 0.1, 5.6, 0.2, 10.6, 0.4, 23.9, 0.9, 14000.3, 2.3, 20370.1, 3.8, 31359.7, 6.4, 53808, 11.4, 100193.7, 22.9, 162747.6, 43.9, 31074, 5.2, 43067.6, 8, 66533.8, 13.5, 99766.8, 21.1, 161156.5, 36.9, 220591.3, 59.5, 32.7, 1.4, 50, 2, 76.5, 3.3, 126.4, 5.4, 241.4, 9.8, 458.1, 16.9, 12.9, 0.5, 19.3, 0.8, 30.9, 1.3, 69.5, 3, 173.1, 7, 420.3, 15.5, 51.7, 2.1, 90.1, 3.7, 159.2, 6.8, 281.9, 12, 507.1, 20.5, 859.5, 31.7)), row.names = c(NA, -120L), class = "data.frame")
尝试的代码
library(ggplot2) df %>% group_by(age) %>% arrange(desc(value)) %>% ggplot(aes(x = age, y = value, fill = disease)) + geom_bar(stat = "identity") + theme(axis.text.x = element_text(angle = 90) )
解决方案
这个需求完全可以实现。你之前的代码未生效的核心原因是:ggplot2中fill参数的堆叠顺序由disease列的全局因子水平决定,分组后的arrange仅改变数据行顺序,不会调整因子水平的排序逻辑,因此无法实现每个年龄组单独排序。
以下是两种可行的实现方法:
方法一:用fct_reorder2动态调整排序
借助forcats包的fct_reorder2函数,可针对每个age组,按value降序重新调整disease的因子顺序:
library(ggplot2) library(forcats) library(dplyr) # 先处理数据:聚合重复的age+disease组合(可选,根据你的数据需求) df_clean <- df %>% group_by(age, disease) %>% summarise(value = sum(value), .groups = "drop") # 绘图 df_clean %>% mutate(disease = fct_reorder2(disease, age, desc(value))) %>% ggplot(aes(x = age, y = value, fill = disease)) + geom_col(position = position_stack(reverse = TRUE)) + theme(axis.text.x = element_text(angle = 90))
说明
fct_reorder2(disease, age, desc(value)):针对每个年龄组,将疾病按value降序排序position_stack(reverse = TRUE):默认堆叠顺序为因子水平升序,设置此参数可让数值大的类别位于条形底部- 数据聚合步骤:你的原始数据存在重复的
age+disease组合,建议先聚合避免重复堆叠
方法二:手动生成分组排序后的复合因子
如果需要更精细的控制,可先为每个年龄组计算疾病排序,再创建包含年龄组的复合因子确保堆叠顺序:
library(ggplot2) library(dplyr) # 聚合数据并生成每个年龄组内的疾病排序 df_sorted <- df %>% group_by(age, disease) %>% summarise(value = sum(value), .groups = "drop") %>% group_by(age) %>% arrange(desc(value), .by_group = TRUE) %>% mutate(disease_sorted = factor(paste(age, disease, sep = "_"), levels = paste(age, disease, sep = "_"))) %>% ungroup() # 绘图 ggplot(df_sorted, aes(x = age, y = value, fill = disease_sorted)) + geom_col() + scale_fill_discrete(labels = function(x) gsub(".*_", "", x)) + theme(axis.text.x = element_text(angle = 90))
说明
- 复合因子
disease_sorted确保每个年龄组内的疾病按value降序排列 scale_fill_discrete将图例标签还原为原始疾病名称,避免显示复合因子的拼接内容
内容的提问来源于stack exchange,提问作者Ghose Bishwajit
相关产品推荐
相关产品推荐

