如何在ggplot中按因子S分组,为列均值绘制横竖参考线?
问题
我有如下格式的tibble数据框:
# A tibble: 6 × 8 N S DG P locus replicate MLE_s_BwS MLE_s_WF <int> <dbl> <int> <int> <int> <int> <dbl> <dbl> 1 100 0 1 1 0 0 -0.174 -0.183 2 100 0 1 1 1 0 0.143 0.143 3 100 0 1 1 2 0 -0.0758 -0.0758 4 100 0.1 10 1 3 0 -0.141 -0.141 5 100 0.1 10 1 4 0 0.102 0.102 6 100 0.1 10 1 5 0 0.102 0.102
我用ggplot分别绘制DG == 1和DG == 10时,MLE_s_BwS与MLE_s_WF的散点图,代码如下:
my_plot_DG1 <- ggplot( estimations_pivoted[estimations_pivoted$DG == 1,], aes(x = MLE_s_WF, y = MLE_s_BwS) ) + geom_point(aes(col = as.factor(S)), alpha = 0.5) + scale_color_manual(values = c("purple2", "yellow")) + geom_abline(linetype = "dotted") + xlim(x_lim) + ylim(y_lim) + theme_minimal()
现在需要按S == 0和S == 0.1分组,分别为MLE_s_BwS的均值绘制水平线、为MLE_s_WF的均值绘制垂直线。目前只能手动计算均值后调用geom_hline,不够优雅,有没有更简洁的实现方式?
解决方案
方法一:用stat_summary自动计算并绘制分组均值线
利用stat_summary可以直接在绘图时按S分组计算均值,自动生成对应水平线和垂直线,无需手动计算:
my_plot_DG1 <- ggplot( estimations_pivoted[estimations_pivoted$DG == 1,], aes(x = MLE_s_WF, y = MLE_s_BwS, color = as.factor(S)) ) + geom_point(alpha = 0.5) + scale_color_manual(values = c("purple2", "yellow")) + # 按S分组绘制MLE_s_WF均值的垂直线 stat_summary( aes(xintercept = after_stat(x)), fun = mean, geom = "vline", linetype = "solid", linewidth = 0.8 ) + # 按S分组绘制MLE_s_BwS均值的水平线 stat_summary( aes(yintercept = after_stat(y)), fun = mean, geom = "hline", linetype = "solid", linewidth = 0.8 ) + geom_abline(linetype = "dotted") + xlim(x_lim) + ylim(y_lim) + theme_minimal()
该方法会自动匹配分组颜色,线条颜色与对应分组的点一致,逻辑简洁高效。
方法二:提前生成汇总数据再绘制
如果需要更灵活控制汇总逻辑,可先通过分组计算得到均值数据,再传入绘图函数:
# 生成按DG、S分组的均值汇总表 summary_data <- estimations_pivoted %>% group_by(DG, S) %>% summarise( mean_BwS = mean(MLE_s_BwS), mean_WF = mean(MLE_s_WF), .groups = "drop" ) # 绘制DG=1的散点图+均值线 my_plot_DG1 <- ggplot( estimations_pivoted[estimations_pivoted$DG == 1,], aes(x = MLE_s_WF, y = MLE_s_BwS, color = as.factor(S)) ) + geom_point(alpha = 0.5) + scale_color_manual(values = c("purple2", "yellow")) + # 匹配DG=1的汇总数据绘制水平线 geom_hline( data = summary_data[summary_data$DG == 1,], aes(yintercept = mean_BwS, color = as.factor(S)), linetype = "solid", linewidth = 0.8 ) + # 匹配DG=1的汇总数据绘制垂直线 geom_vline( data = summary_data[summary_data$DG == 1,], aes(xintercept = mean_WF, color = as.factor(S)), linetype = "solid", linewidth = 0.8 ) + geom_abline(linetype = "dotted") + xlim(x_lim) + ylim(y_lim) + theme_minimal()
这种方式的优势是汇总数据可复用,方便绘制多个不同DG值的图表,也便于查看具体均值数值。
内容的提问来源于stack exchange,提问作者Kiffikiffe
相关产品推荐
相关产品推荐

