如何在ggplot2的geom_bar图中按kiwi数量降序排列Y轴聚类?
按Kiwi计数降序排序聚类柱状图的Y轴
问题描述
我用R语言的ggplot2包绘制了聚类与水果数量的分组柱状图,并用coord_flip()翻转了坐标轴。现在希望将Y轴上的cluster按kiwi的计数从高到低排序,请问该如何实现?
原代码:
df = data.frame() df = data.frame(matrix(df, nrow=200, ncol=2)) colnames(df) <- c("cluster", "name") df$cluster <- sample(20, size = nrow(df), replace = TRUE) df$fruit <- sample(c("banana", "apple", "orange", "kiwi", "plum"), size = nrow(df), replace = TRUE) p = ggplot(df, aes(x = as.factor(cluster), fill = as.factor(fruit)))+ geom_bar(stat = 'count') + theme_classic()+ coord_flip() + theme(axis.text.y = element_text(size = 20), axis.title.x = element_text(size = 20), axis.title.y = element_text(size = 20), axis.text=element_text(size=20)) + theme(legend.text = element_text(size = 20)) + xlab("Cluster")+ ylab("Fruit count") + labs( fill = "") p
当前可视化效果:
解决方案
要实现按kiwi计数排序cluster,核心是将cluster转换为按kiwi计数降序排列的因子。这里提供两种简洁的实现方式:
方法1:使用forcats包的fct_reorder()(推荐)
fct_reorder()可以直接在ggplot的映射中对因子进行排序,无需额外预处理数据:
library(ggplot2) library(forcats) df = data.frame() df = data.frame(matrix(df, nrow=200, ncol=2)) colnames(df) <- c("cluster", "name") df$cluster <- sample(20, size = nrow(df), replace = TRUE) df$fruit <- sample(c("banana", "apple", "orange", "kiwi", "plum"), size = nrow(df), replace = TRUE) # 核心修改:用fct_reorder按kiwi计数降序排列cluster p = ggplot(df, aes(x = fct_reorder(as.factor(cluster), as.numeric(fruit == "kiwi"), .fun = sum, .desc = TRUE), fill = as.factor(fruit)))+ geom_bar(stat = 'count') + theme_classic()+ coord_flip() + theme(axis.text.y = element_text(size = 20), axis.title.x = element_text(size = 20), axis.title.y = element_text(size = 20), axis.text=element_text(size=20)) + theme(legend.text = element_text(size = 20)) + xlab("Cluster")+ ylab("Fruit count") + labs( fill = "") p
参数说明:
fct_reorder()第一个参数:待排序的因子(as.factor(cluster))- 第二个参数:排序依据的数值——
as.numeric(fruit == "kiwi")会将kiwi对应的行转为1,其他水果转为0 .fun = sum:对每个cluster的上述数值求和,得到该cluster的kiwi计数.desc = TRUE:指定按降序排列
方法2:手动计算kiwi计数并排序(无需额外包)
如果不想加载forcats包,可以先计算每个cluster的kiwi数量,再重新设定因子的水平:
library(ggplot2) library(dplyr) df = data.frame() df = data.frame(matrix(df, nrow=200, ncol=2)) colnames(df) <- c("cluster", "name") df$cluster <- sample(20, size = nrow(df), replace = TRUE) df$fruit <- sample(c("banana", "apple", "orange", "kiwi", "plum"), size = nrow(df), replace = TRUE) # 计算每个cluster的kiwi计数 kiwi_counts <- df %>% filter(fruit == "kiwi") %>% count(cluster, name = "kiwi_count") # 合并计数到原数据,无kiwi的cluster计数设为0 df <- df %>% left_join(kiwi_counts, by = "cluster") %>% mutate(kiwi_count = ifelse(is.na(kiwi_count), 0, kiwi_count)) # 按kiwi计数降序设定cluster因子的水平 df$cluster <- factor(df$cluster, levels = df %>% distinct(cluster, kiwi_count) %>% arrange(desc(kiwi_count)) %>% pull(cluster)) # 绘图 p = ggplot(df, aes(x = cluster, fill = as.factor(fruit)))+ geom_bar(stat = 'count') + theme_classic()+ coord_flip() + theme(axis.text.y = element_text(size = 20), axis.title.x = element_text(size = 20), axis.title.y = element_text(size = 20), axis.text=element_text(size=20)) + theme(legend.text = element_text(size = 20)) + xlab("Cluster")+ ylab("Fruit count") + labs( fill = "") p
内容的提问来源于stack exchange,提问作者MeganCole
相关产品推荐
相关产品推荐

