使用dplyr::mutate修改列值时,如何保留嵌套箱线图的分组排序?
嵌套X轴分组顺序错乱问题解决
问题现象
在使用ggh4x创建嵌套箱线图时,先将Tx列设为因子指定顺序(Not Exposed在前,Exposed在后),但后续通过mutate给Tx添加样本量标注后,嵌套X轴的分组顺序出现错误——Exposed跑到了Not Exposed前面。
问题原因
修改Tx内容(添加\n(n = ...))时,原有的因子结构被破坏:新生成的Tx字符串不再是原来的因子水平,R会自动将其转换为字符型。后续生成interaction(Tx, Species)时,会按字符默认的字母顺序排序,而"Exposed"首字母E早于"Not Exposed"的N,导致顺序颠倒。
解决方案
核心思路是保留原始Tx的因子顺序信息,不要直接修改原Tx列,而是通过新建标签列或利用ggh4x的嵌套轴功能来实现。
方法1:手动控制交互项的因子水平
set1 <- structure(list(Tx = c("Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed"), Species = structure(c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L), levels = c("Species1", "Species2"), class = "factor"), Size = c(88.5, 83.3, 59.5, 78, 50.3, 57, 78.2, 59, 85, 59.5, 13.1, 50.1, 55, 60.1, 13.8, 27, 57.1, 53.1, 42, 16, 88.8, 26.2, 62, 108.5, 92.3, 74.4, 77.3, 96, 88.7, 77.8, 50.7, 61.9, 65.1, 63.5, 64, 88.6, 53.8, 82.1, 78.8, 75.6)), row.names = c(NA, -40L), class = c("tbl_df", "tbl", "data.frame")) # 计算分组统计量与样本量 set1 <- set1 %>% mutate( w = mean(Size), max = max(Size), n = n(), .by = c(Tx, Species) ) # 保留原始Tx因子,新建带样本量的标签列 set1 <- set1 %>% mutate( Tx_label = paste0(Tx, '\n(n = ', n, ')'), x_interaction = interaction(Tx, Species, lex.order = TRUE) ) # 手动指定交互项的顺序,确保符合需求 set1$x_interaction <- factor( set1$x_interaction, levels = c( "Not Exposed.Species1", "Not Exposed.Species2", "Exposed.Species1", "Exposed.Species2" ) ) # 绘图 ggplot(set1, aes(x = x_interaction, y = Size)) + geom_boxplot() + geom_jitter(width = 0.1, stroke = 0.5, size = 2) + scale_x_discrete(labels = set1 %>% distinct(x_interaction, Tx_label) %>% pull(Tx_label)) + guides(x = "axis_nested") + theme_classic()
方法2:利用ggh4x的嵌套轴原生功能
这种方式更简洁,不需要手动生成交互项:
set1 <- structure(list(Tx = c("Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Not Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed", "Exposed"), Species = structure(c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L), levels = c("Species1", "Species2"), class = "factor"), Size = c(88.5, 83.3, 59.5, 78, 50.3, 57, 78.2, 59, 85, 59.5, 13.1, 50.1, 55, 60.1, 13.8, 27, 57.1, 53.1, 42, 16, 88.8, 26.2, 62, 108.5, 92.3, 74.4, 77.3, 96, 88.7, 77.8, 50.7, 61.9, 65.1, 63.5, 64, 88.6, 53.8, 82.1, 78.8, 75.6)), row.names = c(NA, -40L), class = c("tbl_df", "tbl", "data.frame")) # 计算统计量与样本量,同时设置Tx的因子顺序 set1 <- set1 %>% mutate( w = mean(Size), max = max(Size), n = n(), .by = c(Tx, Species) ) %>% mutate( Tx = factor(Tx, levels = c("Not Exposed", "Exposed")), Tx_label = paste0(Tx, '\n(n = ', n, ')') ) # 直接用Tx_label作为x轴,通过group指定分组逻辑 ggplot(set1, aes(x = Tx_label, y = Size, group = interaction(Tx, Species))) + geom_boxplot() + geom_jitter(width = 0.1, stroke = 0.5, size = 2) + guides(x = guide_axis_nested(delim = "\n")) + theme_classic()
说明
- 方法1通过手动定义交互项的因子水平,完全控制轴的排列顺序,适合复杂分组场景;
- 方法2利用ggh4x的
guide_axis_nested自动识别标签中的换行符作为分组分隔,代码更简洁,依赖原始Tx的因子顺序来保证分组排列正确。
内容的提问来源于stack exchange,提问作者arnaudm
相关产品推荐
相关产品推荐

