如何在同一ggplot中用双Y轴展示计数与均值
双Y轴ggplot图表解决方案
没问题,我来帮你搞定这个双Y轴的组合图表!首先得明确:ggplot本身对双Y轴的支持是“视觉层面”的,需要手动对齐两个轴的数值范围,同时要处理好分组(year)的对齐问题。下面给你两种可行的方案,选你觉得顺手的就行。
方案一:先预处理数据(更直观)
先把需要的统计量(集群计数、divergence均值)提前计算好,这样绘图逻辑更清晰,也方便调整轴的比例。
步骤1:整理数据集
你给的数据集格式有点乱,先转换成R能识别的格式:
df <- data.frame( divergence = c(0.34, 0.89, 1.22, 1.11, 0.67, 0.89, 1.12, 0.4, 0.67, 0.89, 0.56, 1.22, 1.12, 0.4, 0.67, 0.89, 0.11, 1.33, 1.11, 1, 0.89, 1, 0.45, 0.23, 0.89, 0.8, 0.6, 1.11, 0.34, 0.78, 2.1, 0.89, 0.89, 0.34, 1.55, 1.11, 1.11, 1, 0.34, 0.67, 0.56, 1, 0.34, 0.67, 0, 0.67, 0.67, 0.34, 0.45), year = factor(c(2015,2015,2015,2015,2015,2015,2015,2015,2015,2015,2015,2015,2015,2015,2016,2016,2016,2016,2016,2016,2016,2016,2016,2016,2016,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017)), cluster = factor(c("A","A","A","B","B","B","B","B","B","B","B","B","B","B","A","A","B","C","C","C","C","C","C","C","C","A","A","A","B","B","C","C","C","C","A","A","A","A","A","A","A","B","C","B","B","B","B","B","B")) )
步骤2:预处理统计量
用dplyr分组计算每个cluster+year的计数和divergence均值:
library(dplyr) summary_df <- df %>% group_by(cluster, year) %>% summarise( count = n(), # 集群计数 mean_div = mean(divergence, na.rm = TRUE) # divergence均值 ) %>% ungroup()
步骤3:绘制双Y轴图表
计算转换因子,让均值的数值范围匹配计数的范围,然后组合图层:
library(ggplot2) # 计算轴的转换因子(让两个变量的最大值对齐) max_count <- max(summary_df$count) max_mean <- max(summary_df$mean_div) scale_factor <- max_count / max_mean ggplot(summary_df, aes(x = cluster)) + # 主Y轴:分组柱状图(集群计数) geom_bar(aes(y = count, fill = year), stat = "identity", position = "dodge", alpha = 0.7) + # 次Y轴:均值的点和线 geom_point(aes(y = mean_div * scale_factor, color = year), size = 3, position = position_dodge(width = 0.9)) + geom_line(aes(y = mean_div * scale_factor, color = year, group = year), position = position_dodge(width = 0.9)) + # 设置双Y轴 scale_y_continuous( name = "Cluster Count", sec.axis = sec_axis(~./scale_factor, name = "Mean Divergence") ) + # 美化颜色和主题 scale_fill_brewer(palette = "Blues") + scale_color_brewer(palette = "Reds") + labs(title = "Cluster Count vs Mean Divergence by Year", x = "Cluster") + theme_minimal() + theme( legend.position = "top", plot.title = element_text(hjust = 0.5) )
方案二:直接用ggplot的stat函数(更简洁)
如果不想预处理数据,可以直接用stat_count和stat_summary来计算统计量,省去中间步骤:
library(ggplot2) library(dplyr) # 先计算转换因子 count_max <- df %>% group_by(cluster, year) %>% summarise(n=n()) %>% pull(n) %>% max() mean_max <- df %>% group_by(cluster, year) %>% summarise(m=mean(divergence)) %>% pull(m) %>% max() scale_factor <- count_max / mean_max ggplot(df, aes(x = cluster)) + # 主Y轴:计数柱状图 geom_bar(aes(fill = year), stat = "count", position = "dodge", alpha = 0.7) + # 次Y轴:均值的点和线 stat_summary(aes(y = divergence * scale_factor, color = year), fun = mean, geom = "point", size = 3, position = position_dodge(width = 0.9)) + stat_summary(aes(y = divergence * scale_factor, color = year, group = year), fun = mean, geom = "line", position = position_dodge(width = 0.9)) + # 设置双Y轴 scale_y_continuous( name = "Cluster Count", sec.axis = sec_axis(~./scale_factor, name = "Mean Divergence") ) + # 美化设置 scale_fill_brewer(palette = "Blues") + scale_color_brewer(palette = "Reds") + labs(title = "Cluster Count vs Mean Divergence by Year", x = "Cluster") + theme_minimal() + theme(legend.position = "top")
注意事项
- 双Y轴容易造成视觉误导,一定要确保两个变量的对比是有意义的,并且转换因子是明确的(这里用最大值比例来对齐,是比较合理的方式)。
position_dodge(width = 0.9)是为了让点/线和柱状图的分组位置对齐,0.9是geom_bar默认的dodge宽度,保持一致就能对齐。
内容的提问来源于stack exchange,提问作者Nido
相关产品推荐
相关产品推荐

