如何在R分组柱状图中调整变量标签格式并按链接方法分组?
问题需求
现有如下data.table格式的数据:
dataset variable ARI 1: pcaZ0 pearson, single 0.6984690 2: pcaZ0 pearson, complete 0.6984690 3: pcaZ0 pearson, average 0.6984690 4: pcaZ0 spearman, single 0.6984690 5: pcaZ0 spearman, complete 0.6984690 6: pcaZ0 spearman, average 0.6984690 7: pcaZ0 cosine, single 0.7238611 8: pcaZ0 cosine, complete 0.7238611 9: pcaZ0 cosine, average 0.7109783 10: pcaZ0 euclidean, single 0.5177371 11: pcaZ0 euclidean, complete 0.5177371 12: pcaZ0 euclidean, average 0.5177371 13: pcaZ1 pearson, single 0.5429425 14: pcaZ1 pearson, complete 0.9619119 15: pcaZ1 pearson, average 0.5429425 16: pcaZ1 spearman, single 0.5317401 17: pcaZ1 spearman, complete 0.5317401 18: pcaZ1 spearman, average 0.7371173 19: pcaZ1 cosine, single 0.5220314 20: pcaZ1 cosine, complete 0.9434279 21: pcaZ1 cosine, average 0.8089993 22: pcaZ1 euclidean, single 0.5177371 23: pcaZ1 euclidean, complete 0.5177371 24: pcaZ1 euclidean, average 0.5177371 25: modpcaZ0 pearson, single 0.5251420 26: modpcaZ0 pearson, complete 0.8485167 27: modpcaZ0 pearson, average 0.9596045 28: modpcaZ0 spearman, single 0.5251420 29: modpcaZ0 spearman, complete 0.8485167 30: modpcaZ0 spearman, average 0.8838628 31: modpcaZ0 cosine, single 0.5083105 32: modpcaZ0 cosine, complete 0.9596045 33: modpcaZ0 cosine, average 0.9596045 34: modpcaZ0 euclidean, single 0.5203030 35: modpcaZ0 euclidean, complete 0.6717862 36: modpcaZ0 euclidean, average 0.5360825 37: modpcaZ1 pearson, single 0.5360825 38: modpcaZ1 pearson, complete 0.8485167 39: modpcaZ1 pearson, average 0.8838628 40: modpcaZ1 spearman, single 0.5630128 41: modpcaZ1 spearman, complete 0.8485167 42: modpcaZ1 spearman, average 0.8485167 43: modpcaZ1 cosine, single 0.5360825 44: modpcaZ1 cosine, complete 0.8314749 45: modpcaZ1 cosine, average 0.9400379 46: modpcaZ1 euclidean, single 0.5360825 47: modpcaZ1 euclidean, complete 0.7239638 48: modpcaZ1 euclidean, average 0.5487061
使用以下代码生成分组柱状图:
library(tidyverse) library(ggplot2) library(gridExtra) library(ggtext) library(khroma) # grouped bar charts ggplot(b2, aes(y=variable, x=ARI, fill=dataset)) + geom_col(position=position_dodge(), width=0.6) + scale_fill_manual(name=NULL, breaks=c("pcaZ0", "pcaZ1", "modpcaZ0", "modpcaZ1"), labels=c("non-standard.", "standard.", "non-standard.<br>*without 71-76*", "standard.<br>*without 71-76*"), values=c(mypal2)) + scale_x_continuous(expand=c(0, 0)) + labs(y=NULL, x="adjusted Rand index") + theme_classic() + theme(axis.text.x=element_markdown(), legend.text=element_markdown(), legend.position=c(0.9, 0.9), panel.grid.major.x=element_line(color="lightgray", size=0.25))
其中mypal2的定义为:
mypal <- colour("okabeito")(8) mypal <- mypal[c(2:8, 1)] names(mypal) <- NULL mypal2 <- mypal[-c(2, 4, 6)] palette(mypal2)
需要实现两个优化:
- 将变量按**链接方法(single/complete/average)**分组展示
- 把
variable列中如pearson, single的标签,改为第一行显示相似性度量(如pearson)、第二行显示斜体链接方法(如*single*)的格式
解决方法
要实现需求,需先处理数据结构,再修改绘图代码,具体步骤如下:
1. 数据预处理
拆分variable列,生成带格式的新标签,并按链接方法排序:
用data.table语法处理:
library(data.table) # 确保数据为data.table格式 setDT(b2) # 拆分variable列为度量和链接方法两列 b2[, c("metric", "linkage") := tstrsplit(variable, ", ", fixed=TRUE)] # 生成带换行和斜体的新标签 b2[, new_variable := paste0(metric, "<br>*", linkage, "*")] # 按链接方法排序,保证同组变量相邻 b2[, linkage := factor(linkage, levels=c("single", "complete", "average"))] b2 <- b2[order(linkage, metric)] # 将新标签转为因子,锁定排序顺序 b2[, new_variable := factor(new_variable, levels=unique(new_variable))]
用tidyverse语法处理:
b2 <- b2 %>% separate(variable, into=c("metric", "linkage"), sep=", ", remove=FALSE) %>% mutate( linkage = factor(linkage, levels=c("single", "complete", "average")), new_variable = paste0(metric, "<br>*", linkage, "*") ) %>% arrange(linkage, metric) %>% mutate(new_variable = factor(new_variable, levels=unique(new_variable)))
2. 修改绘图代码
替换y轴变量为新标签,启用Markdown解析,同时添加分组面板:
ggplot(b2, aes(y=new_variable, x=ARI, fill=dataset)) + geom_col(position=position_dodge(), width=0.6) + # 按链接方法纵向分组,每组独立显示y轴 facet_grid(linkage ~ ., scales="free_y", space="free_y") + scale_fill_manual(name=NULL, breaks=c("pcaZ0", "pcaZ1", "modpcaZ0", "modpcaZ1"), labels=c("非标准化", "标准化", "非标准化<br>*不含71-76*", "标准化<br>*不含71-76*"), values=mypal2) + scale_x_continuous(expand=c(0, 0)) + labs(y=NULL, x="调整兰德指数") + theme_classic() + theme( axis.text.y=element_markdown(), # 解析y轴的Markdown格式标签 legend.text=element_markdown(), legend.position=c(0.9, 0.9), panel.grid.major.x=element_line(color="lightgray", size=0.25), strip.background=element_blank(), # 去掉分组面板的灰色背景 strip.text.y=element_text(angle=0, hjust=0) # 调整分组标签的显示角度 )
可选调整:
如果不需要分面板展示,仅需同链接方法的变量相邻,去掉facet_grid相关代码即可,此时new_variable的因子顺序已保证分组效果。
内容的提问来源于stack exchange,提问作者wantingtoimprove
相关产品推荐
相关产品推荐

