You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在同一ggplot中用双Y轴展示计数与均值

双Y轴ggplot图表解决方案

没问题,我来帮你搞定这个双Y轴的组合图表!首先得明确:ggplot本身对双Y轴的支持是“视觉层面”的,需要手动对齐两个轴的数值范围,同时要处理好分组(year)的对齐问题。下面给你两种可行的方案,选你觉得顺手的就行。

方案一:先预处理数据(更直观)

先把需要的统计量(集群计数、divergence均值)提前计算好,这样绘图逻辑更清晰,也方便调整轴的比例。

步骤1:整理数据集

你给的数据集格式有点乱,先转换成R能识别的格式:

df <- data.frame(
  divergence = c(0.34, 0.89, 1.22, 1.11, 0.67, 0.89, 1.12, 0.4, 0.67, 0.89, 0.56, 1.22, 1.12, 0.4, 0.67, 0.89, 0.11, 1.33, 1.11, 1, 0.89, 1, 0.45, 0.23, 0.89, 0.8, 0.6, 1.11, 0.34, 0.78, 2.1, 0.89, 0.89, 0.34, 1.55, 1.11, 1.11, 1, 0.34, 0.67, 0.56, 1, 0.34, 0.67, 0, 0.67, 0.67, 0.34, 0.45),
  year = factor(c(2015,2015,2015,2015,2015,2015,2015,2015,2015,2015,2015,2015,2015,2015,2016,2016,2016,2016,2016,2016,2016,2016,2016,2016,2016,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017,2017)),
  cluster = factor(c("A","A","A","B","B","B","B","B","B","B","B","B","B","B","A","A","B","C","C","C","C","C","C","C","C","A","A","A","B","B","C","C","C","C","A","A","A","A","A","A","A","B","C","B","B","B","B","B","B"))
)

步骤2:预处理统计量

用dplyr分组计算每个cluster+year的计数和divergence均值:

library(dplyr)

summary_df <- df %>%
  group_by(cluster, year) %>%
  summarise(
    count = n(),  # 集群计数
    mean_div = mean(divergence, na.rm = TRUE)  # divergence均值
  ) %>%
  ungroup()

步骤3:绘制双Y轴图表

计算转换因子,让均值的数值范围匹配计数的范围,然后组合图层:

library(ggplot2)

# 计算轴的转换因子(让两个变量的最大值对齐)
max_count <- max(summary_df$count)
max_mean <- max(summary_df$mean_div)
scale_factor <- max_count / max_mean

ggplot(summary_df, aes(x = cluster)) +
  # 主Y轴:分组柱状图(集群计数)
  geom_bar(aes(y = count, fill = year), stat = "identity", position = "dodge", alpha = 0.7) +
  # 次Y轴:均值的点和线
  geom_point(aes(y = mean_div * scale_factor, color = year), size = 3, position = position_dodge(width = 0.9)) +
  geom_line(aes(y = mean_div * scale_factor, color = year, group = year), position = position_dodge(width = 0.9)) +
  # 设置双Y轴
  scale_y_continuous(
    name = "Cluster Count",
    sec.axis = sec_axis(~./scale_factor, name = "Mean Divergence")
  ) +
  # 美化颜色和主题
  scale_fill_brewer(palette = "Blues") +
  scale_color_brewer(palette = "Reds") +
  labs(title = "Cluster Count vs Mean Divergence by Year", x = "Cluster") +
  theme_minimal() +
  theme(
    legend.position = "top",
    plot.title = element_text(hjust = 0.5)
  )

方案二:直接用ggplot的stat函数(更简洁)

如果不想预处理数据,可以直接用stat_count和stat_summary来计算统计量,省去中间步骤:

library(ggplot2)
library(dplyr)

# 先计算转换因子
count_max <- df %>% group_by(cluster, year) %>% summarise(n=n()) %>% pull(n) %>% max()
mean_max <- df %>% group_by(cluster, year) %>% summarise(m=mean(divergence)) %>% pull(m) %>% max()
scale_factor <- count_max / mean_max

ggplot(df, aes(x = cluster)) +
  # 主Y轴:计数柱状图
  geom_bar(aes(fill = year), stat = "count", position = "dodge", alpha = 0.7) +
  # 次Y轴:均值的点和线
  stat_summary(aes(y = divergence * scale_factor, color = year), 
               fun = mean, geom = "point", size = 3, position = position_dodge(width = 0.9)) +
  stat_summary(aes(y = divergence * scale_factor, color = year, group = year), 
               fun = mean, geom = "line", position = position_dodge(width = 0.9)) +
  # 设置双Y轴
  scale_y_continuous(
    name = "Cluster Count",
    sec.axis = sec_axis(~./scale_factor, name = "Mean Divergence")
  ) +
  # 美化设置
  scale_fill_brewer(palette = "Blues") +
  scale_color_brewer(palette = "Reds") +
  labs(title = "Cluster Count vs Mean Divergence by Year", x = "Cluster") +
  theme_minimal() +
  theme(legend.position = "top")

注意事项

  • 双Y轴容易造成视觉误导,一定要确保两个变量的对比是有意义的,并且转换因子是明确的(这里用最大值比例来对齐,是比较合理的方式)。
  • position_dodge(width = 0.9)是为了让点/线和柱状图的分组位置对齐,0.9是geom_bar默认的dodge宽度,保持一致就能对齐。

内容的提问来源于stack exchange,提问作者Nido

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:47:52