ggplot2挑战:在多阶段蛋白质累积丰度图中标注特定蛋白
问题:在蛋白质累积丰度图中标注特定蛋白
我们需要在展示不同细胞发育阶段蛋白质累积丰度的图表中标注特定蛋白(如SYMBOL为'ELANE'的蛋白),以展示其在蛋白质组排名中的升降变化。当前已绘制出有效图表,但尝试的标注代码无效。
原绘图代码
dat.pro.mols %>% ggplot() + geom_point(data = . %>% filter(stage == 'MB'), aes(x= 1:3156, y= cum_Mol), color = "#0000e3") + geom_point(data = . %>% filter(stage == 'PM'), aes(x= 1:3156, y= cum_Mol), color = "#a043ec") + geom_point(data = . %>% filter(stage == 'MC'), aes(x= 1:3156, y= cum_Mol), color = "#a80092") + geom_point(data = . %>% filter(stage == 'MM'), aes(x= 1:3156, y= cum_Mol), color = "#ca0068") + geom_point(data = . %>% filter(stage == 'B'), aes(x= 1:3156, y= cum_Mol), color = "#ff7763") + geom_point(data = . %>% filter(stage == 'PMN'), aes(x= 1:3156, y= cum_Mol), color = "#c8a600") + geom_hline(yintercept = 0.5, linetype="dashed",alpha=0.7) + xlab("Protein rank by abundance") + ylab("Cumulative protein abundance") + theme(legend.position = "none") + xlim(0,70)
无效的标注代码
geom_text(data = . %>% filter(SYMBOL == 'ELANE'), aes(y= cum_Mol, x = cum_Mol, label = SYMBOL)) +
问题分析
- 原代码中x轴用
1:3156代表蛋白丰度排名,但标注时错误将x映射为cum_Mol,完全偏离x轴含义,导致标注位置错误。 - 每个细胞阶段的蛋白排名独立,需先为每个阶段计算对应蛋白的排名位置,才能准确定位标注。
解决方案
步骤1:预处理数据,计算每个阶段的蛋白排名
先确保数据按每个阶段的蛋白丰度降序排列,再为每个阶段生成排名:
# 若数据未按丰度降序排列,需先添加arrange步骤 dat.pro.mols <- dat.pro.mols %>% group_by(stage) %>% # 未排序时添加:arrange(desc(median_quant)) %>% mutate(rank = row_number()) %>% ungroup()
步骤2:修正绘图代码,添加正确标注
替换原绘图代码中x轴的1:3156为计算好的rank,并添加正确的geom_text标注:
dat.pro.mols %>% ggplot() + # 替换x轴映射为计算好的rank geom_point(data = . %>% filter(stage == 'MB'), aes(x= rank, y= cum_Mol), color = "#0000e3") + geom_point(data = . %>% filter(stage == 'PM'), aes(x= rank, y= cum_Mol), color = "#a043ec") + geom_point(data = . %>% filter(stage == 'MC'), aes(x= rank, y= cum_Mol), color = "#a80092") + geom_point(data = . %>% filter(stage == 'MM'), aes(x= rank, y= cum_Mol), color = "#ca0068") + geom_point(data = . %>% filter(stage == 'B'), aes(x= rank, y= cum_Mol), color = "#ff7763") + geom_point(data = . %>% filter(stage == 'PMN'), aes(x= rank, y= cum_Mol), color = "#c8a600") + # 添加正确标注:x用rank,y用cum_Mol,调整位置避免重叠 geom_text(data = . %>% filter(SYMBOL == 'ELANE'), aes(x= rank, y= cum_Mol, label = SYMBOL, color = stage), hjust = -0.1, vjust = 0.5, size = 4) + geom_hline(yintercept = 0.5, linetype="dashed",alpha=0.7) + xlab("Protein rank by abundance") + ylab("Cumulative protein abundance") + theme(legend.position = "none") + xlim(0,70)
代码说明
mutate(rank = row_number()):为每个阶段的蛋白生成从1开始的排名,确保x轴排名与数据对应。geom_text中hjust = -0.1:将文字向右偏移,避免与数据点重叠;color = stage让标注文字颜色匹配对应阶段的点颜色,便于区分。- 若数据未按丰度降序排列,需在
mutate前添加arrange(desc(median_quant)),保证排名按丰度从高到低排序。
内容的提问来源于stack exchange,提问作者Sebastian Hesse
相关产品推荐
相关产品推荐

