使用ggplot2在分组柱状图上叠加散点与误差条的技术问题
解决ggplot2柱状图叠加散点与误差条的三个问题
我来帮你逐一搞定这几个问题,咱们先理清楚问题根源,再给出完整的修正方案:
1. 调整rep散点的正确顺序
你遇到的散点顺序混乱、和柱状图不对齐的问题,主要是两个原因:
- 原始散点的
aes里没关联scientist分组,导致position_dodge无法和柱状图的分组匹配 rep默认是字符型,ggplot会按字母排序,但我们需要它按rep1→rep2→rep3的逻辑顺序排列
修正方法很简单:把rep转成有序因子,同时给散点的aes加上group=scientist,让散点跟着柱状图的分组进行偏移对齐。
2. 将柱状图改为显示中位数(或均值)
原来用stat="identity"直接绘制原始数据,会把同一分组的所有值叠加起来,所以看起来像是最大值。我们需要先对数据做汇总,计算每个cell+scientist+timepoint分组的中位数(或均值),再用汇总后的数据画柱状图。
3. 添加包含最值的误差条
同样在数据汇总时,顺便算出每个分组的最小值和最大值,然后用geom_errorbar绘制,注意要设置和柱状图一致的position_dodge,保证误差条和柱子完全对齐。
完整修正代码
先加载需要的包:
library(ggplot2) library(dplyr) library(RColorBrewer)
生成数据并做预处理与汇总:
set.seed(123) # 设置随机种子让结果可复现 mydf <- data.frame( cell=paste0("cell", rep(1:3, each=12)), scientist=paste0("scientist", rep(rep(rep(1:2, each=3), 2), 3)), timepoint=paste0("time", rep(rep(1:2, each=6), 3)), rep=paste0("rep", rep(1:3, 12)), value=runif(36)*100 ) # 把rep转为有序因子,强制指定顺序 mydf$rep <- factor(mydf$rep, levels = paste0("rep", 1:3), ordered = TRUE) # 汇总数据:计算每个分组的中位数、最小值、最大值 summary_df <- mydf %>% group_by(cell, scientist, timepoint) %>% summarise( median_val = median(value), min_val = min(value), max_val = max(value), .groups = "drop" )
最后绘制图形:
myPal <- brewer.pal(3, "Set2")[1:2] myPal2 <- brewer.pal(3, "Set1") outfile <- "test.pdf" pdf(file=outfile, height=10, width=10) ggplot() + # 用汇总数据画柱状图,按scientist分组偏移 geom_bar(data=summary_df, aes(cell, median_val, fill=scientist), stat="identity", position=position_dodge(.9), width=0.8) + # 添加误差条,对应最值,偏移量和柱状图保持一致 geom_errorbar(data=summary_df, aes(cell, ymin=min_val, ymax=max_val, group=scientist), position=position_dodge(.9), width=0.2) + # 用原始数据画散点,group=scientist确保和柱子对齐,color区分重复样本 geom_point(data=mydf, aes(cell, value, color=rep, group=scientist), position=position_dodge(.9), size=5) + facet_grid(timepoint~., scales="free_x", space="free_x") + scale_y_continuous("% of total cells") + scale_fill_manual(values=myPal) + scale_color_manual(values=myPal2) + labs(x = "Cell Type", y = "% of total cells") + theme_bw() dev.off()
关键修改说明:
- 分开指定图层的数据:柱状图和误差条用汇总后的
summary_df,散点用原始数据 - 所有需要偏移对齐的图层都加上
group=scientist,并使用相同的position_dodge(.9) - 把
rep转为有序因子,强制散点按我们想要的顺序排列 - 如果想把柱状图改成显示均值,只需要把汇总里的
median_val换成mean_val,用mean(value)计算即可
内容的提问来源于stack exchange,提问作者DaniCee
相关产品推荐
相关产品推荐

