基于GOA与depthcenter排序ggplot2气泡图X轴Sample变量
问题:ggplot2气泡图坐标轴批量自定义排序
我有一个超10000行的数据集,已经用ggplot2绘制了气泡图,但需要对坐标轴做自定义排序。我知道可以用factor()结合levels列表来排序变量,但变量数量太多(有数百个基因注释变量),手动编写levels完全不现实。
我之前使用factor()的示例:
primary_order_list4 <- c("RHM","TS","SCS","LIL","STN","STS") fenorm.gene.count$Pond <- factor(fenorm.gene.count$Pond,levels=primary_order_list4)
X轴排序需求
Sample列格式为「地点+年份+深度」(例如SCS19-3),排序规则:
- 先按对应池塘的
GOA值升序排序(比如GOA为6.4的RHM要排在GOA为18.7的SCS之前) - 同一池塘内,再按
depthcenter从浅(0.5)到深(19.5)排序
我尝试过先按depthcenter和GOA升序排序数据框,但ggplot绘制出的X轴仍为字母顺序:
fenorm.gene.count <- fenorm.gene.count[order(fenorm.gene.count$depthcenter),] ugh <- fenorm.gene.count[order(fenorm.gene.count$GOA),] ughplot <- ggplot(fenorm.gene.count, aes(x=Sample,y=HMM,fill=Pond)) + geom_point(aes(size=ifelse(total_count==0, NA, total_count)))
另外Y轴(HMM列,基因注释)也需要类似的自定义排序逻辑。
数据集示例
dput(head(fenorm.gene.count)) structure(list(Project = c("Ga0598240", "Ga0598240", "Ga0598240", "Ga0598240", "Ga0598240", "Ga0598240"), category = c("iron_reduction", "iron_reduction", "iron_oxidation", "iron_oxidation", "iron_reduction", "iron_oxidation"), HMM = c("MtrC_TIGR03507", "OmcS", "OmcF", "Cyc2_repCluster2", "PF01032-FecCD-YfhA-FpvE-YfeCD_transport_family", "PvuCD-FhuB-CbrBC-FeuB-CbrC-YfhA"), total_count = c(10L, 1L, 17L, 55L, 15L, 1L), norm = c(0.00136836343732895, 0.000136836343732895, 0.00232621784345922, 0.00752599890530925, 0.00205254515599343, 0.000136836343732895), Sample = c("LIL19-1", "LIL19-1", "LIL19-1", "LIL19-1", "LIL19-1", "LIL19-1"), Portal = c(3300064049, 3300064049, 3300064049, 3300064049, 3300064049, 3300064049), G2 = c("N", "N", "N", "N", "N", "N"), Pond = c("LIL", "LIL", "LIL", "LIL", "LIL", "LIL"), year = c(2019L, 2019L, 2019L, 2019L, 2019L, 2019L ), depthcenter = c(0.5, 0.5, 0.5, 0.5, 0.5, 0.5), toplow = c("top", "top", "top", "top", "top", "top"), GOA = c(19.4, 19.4, 19.4, 19.4, 19.4, 19.4), elevation = c(8.2, 8.2, 8.2, 8.2, 8.2, 8.2 )), row.names = 65:70, class = "data.frame")
解决方案
1. X轴(Sample)排序
核心思路是先基于GOA和depthcenter自动生成排序规则,再将Sample转为带对应levels的因子,无需手动编写levels:
# 提取每个Sample对应的唯一GOA和depthcenter值(同一Sample的这两个值一致) sample_order_df <- fenorm.gene.count %>% distinct(Sample, Pond, GOA, depthcenter) %>% arrange(GOA, depthcenter) # 按GOA升序、再按depthcenter升序排列 # 将Sample转为因子,levels取排序后的Sample值 fenorm.gene.count$Sample <- factor(fenorm.gene.count$Sample, levels = sample_order_df$Sample)
2. Y轴(HMM)排序
如果Y轴需要按分组或统计指标排序(比如按category分组后,再按total_count均值降序),用同样逻辑自动生成levels:
# 按category分组,计算每个HMM的total_count均值并排序 hmm_order_df <- fenorm.gene.count %>% group_by(category, HMM) %>% summarise(mean_count = mean(total_count), .groups = "drop") %>% arrange(category, desc(mean_count)) # 先按category排序,再按均值降序 # 将HMM转为因子,levels取排序后的HMM值 fenorm.gene.count$HMM <- factor(fenorm.gene.count$HMM, levels = hmm_order_df$HMM)
3. 重新绘制气泡图
现在绘图时,坐标轴会自动按自定义规则排序:
ughplot <- ggplot(fenorm.gene.count, aes(x=Sample, y=HMM, fill=Pond)) + geom_point(aes(size=ifelse(total_count==0, NA, total_count))) + # 若X轴标签拥挤,可添加旋转优化显示 theme(axis.text.x = element_text(angle = 45, hjust = 1))
关键说明
- 仅排序数据框不会改变ggplot的坐标轴顺序,因为ggplot默认将字符型变量视为无序因子,按字母排序。必须将变量转为带自定义levels的因子才能生效。
- 用
dplyr的distinct()和arrange()可自动生成排序后的levels列表,完美适配海量变量场景,无需手动输入。
内容的提问来源于stack exchange,提问作者Geomicro
相关产品推荐
相关产品推荐

