ggplot中值相同时Y轴标签按字母顺序排列的实现方法
解决ggplot2中离散Y轴的双重排序问题
嘿,我明白你遇到的困扰了——既要按中位数排序,又要在中位数相同时让类别按字母顺序排列,还得和样本量注释的顺序对齐对吧?我来给你个靠谱的解决方案:
问题回顾
你想要实现两个核心目标:
- Y轴按
sal的中位数从低到高排序 - 当多个类别的中位数相同时,Y轴标签按字母顺序排列(而非默认的反字母序),这样才能和你添加的样本量注释顺序匹配
你的示例里E和D中位数相同,但当前Y轴顺序是A、B、E、D、C,期望的是A、B、D、E、C,和pivot_df的排序一致。
核心解决方案
与其依赖reorder的默认行为,不如手动提前生成符合要求的排序规则,再把这个规则传给ggplot,这样就能完全掌控排序逻辑了。
步骤1:生成自定义排序的类别顺序
先在pivot_df的arrange环节多加上species,这样当中位数相同时,会自动按字母升序排列:
set.seed(123) df <- data.frame( species = LETTERS[seq(from = 1, to = 5)], sal = round(rnorm(n = 5, mean = 27, sd = .01), 2), num = sample(x = 1:20, size = 20, replace = F) ) # 关键修改:arrange时先按median_sal,再按species字母序 pivot_df <- df %>% group_by(species) %>% summarize(n = n(), median_sal = median(sal, na.rm = T)) %>% arrange(median_sal, species)
步骤2:在ggplot中使用预定义的顺序
把Y轴的映射从reorder(species, -sal, FUN = median)改成factor(species, levels = pivot_df$species),这样就能严格遵循我们提前生成的顺序:
ggplot( data = subset(df, !is.na(sal)), aes(y = factor(species, levels = pivot_df$species), x = sal) ) + geom_boxplot(outlier.shape = 1, outlier.size = 1, orientation = "y") + coord_cartesian(clip = "off") + annotation_custom(grid::textGrob(pivot_df$n, x = 1.035, y = c(0.89, 0.70, 0.51, 0.32, 0.13), gp = grid::gpar(cex = 0.6) )) + annotation_custom(grid::textGrob(expression(bold(underline("N"))), x = 1.035, y = 1.02, gp = grid::gpar(cex = 0.7) )) + ylab("") + xlab("") + theme( axis.text.y = element_text(size = 7, face = "italic"), axis.text.x = element_text(size = 7), axis.title.x = element_text(size = 9, face = "bold"), axis.line = element_line(colour = "black"), panel.background = element_blank(), panel.grid.minor = element_blank(), panel.border = element_rect(colour = "black", fill = NA, size = 1), panel.grid.major = element_line(colour = "#E0E0E0"), plot.title = element_text(hjust = 0.5), plot.margin = margin(21, 40, 20, 20) )
为什么这个方法好用?
reorder函数在遇到相同排序值时,会按原始数据的出现顺序或默认因子顺序排列,这通常不是我们想要的字母序- 手动定义
levels的方式,让我们可以完全控制排序优先级:先按中位数升序,再按类别字母升序,完美匹配你的需求,同时和样本量注释的顺序也能完全对齐
内容的提问来源于stack exchange,提问作者Nate
相关产品推荐
相关产品推荐

