You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在ggplot中连接嵌套分组内各series的中位数?

解决ggplot分组中位数连线问题

问题场景

我需要绘制多分组箱线图,并在每个(setting, sex)组合内,将series(A、B、C)的中位数用线条连接。原始绘图代码如下:

n <- 60
data <- data.frame(series=rep(LETTERS[1:3], n/3), 
                   sex=rep(c("F","M"),each=30), 
                   setting=rep(c("wild","rural"),n/2),
                   fit=rnorm(n))

ggplot(data,aes(x=sex, y=fit, fill=series)) +
geom_boxplot(width=.3,aes( alpha=.5,color=sex),
          lwd=0.8, position = position_dodge(width = 0.6)) + 
facet_grid(~setting) +
stat_summary(fun.y=median, geom="point", shape=23, size=2,
           position=position_dodge(width = 0.6)) +
geom_text(aes(y=-2.5, label=series),  position=position_dodge(width=0.6)) +
geom_point(shape=20,alpha=0.2,position=position_jitterdodge(dodge.width = 0.6,jitter.width = 0.25))+
theme_blank()

尝试用stat_summary+geom_line连线时,要么线条无法和箱线图对齐,要么报错:Error in geom_line():Problem while computing aesthetics.ℹ Error occurred in the 4th layer. Caused by error in FUN():! object 'series' not found。

两种可行解决方案

方案1:调整stat_summary的分组映射

核心是让连线的分组逻辑和箱线图保持一致,x轴沿用sex,同时指定group=series,并使用相同的position_dodge参数对齐:

n <- 60
data <- data.frame(series=rep(LETTERS[1:3], n/3), 
                   sex=rep(c("F","M"),each=30), 
                   setting=rep(c("wild","rural"),n/2),
                   fit=rnorm(n))

ggplot(data,aes(x=sex, y=fit, fill=series)) +
  geom_boxplot(width=.3,aes(alpha=.5, color=sex),
               lwd=0.8, position = position_dodge(width = 0.6)) + 
  facet_grid(~setting) +
  # 保留中位数点
  stat_summary(fun.y=median, geom="point", shape=23, size=2,
               position=position_dodge(width = 0.6)) +
  # 新增中位数连线:关键是group=series,匹配箱线图的dodge宽度
  stat_summary(fun.y=median, geom="line", color="red",
               aes(group=series),
               position=position_dodge(width = 0.6)) +
  geom_text(aes(y=-2.5, label=series),  position=position_dodge(width=0.6)) +
  geom_point(shape=20,alpha=0.2,position=position_jitterdodge(dodge.width = 0.6,jitter.width = 0.25))+
  theme_bw()  # 替换不存在的theme_blank()为标准主题

方案2:提前计算中位数(更直观可控)

先聚合出每个分组的中位数,再用geom_line绘制,避免stat_summary的分组歧义:

library(dplyr)

n <- 60
data <- data.frame(series=rep(LETTERS[1:3], n/3), 
                   sex=rep(c("F","M"),each=30), 
                   setting=rep(c("wild","rural"),n/2),
                   fit=rnorm(n))

# 提前计算各分组中位数
median_df <- data %>%
  group_by(setting, sex, series) %>%
  summarise(median_fit = median(fit), .groups = "drop")

ggplot(data,aes(x=sex, y=fit, fill=series)) +
  geom_boxplot(width=.3,aes(alpha=.5, color=sex),
               lwd=0.8, position = position_dodge(width = 0.6)) + 
  facet_grid(~setting) +
  stat_summary(fun.y=median, geom="point", shape=23, size=2,
               position=position_dodge(width = 0.6)) +
  # 使用预计算的中位数数据绘制连线
  geom_line(data=median_df, aes(x=sex, y=median_fit, group=series),
            color="red", position=position_dodge(width = 0.6)) +
  geom_text(aes(y=-2.5, label=series),  position=position_dodge(width=0.6)) +
  geom_point(shape=20,alpha=0.2,position=position_jitterdodge(dodge.width = 0.6,jitter.width = 0.25))+
  theme_bw()

关键说明

  • group=series:确保在每个(setting, sex)子组内,按series连接对应的中位数点。
  • position_dodge(width = 0.6):和箱线图的dodge宽度完全一致,保证线条与箱线图的中位数点精准对齐。
  • 替换theme_blank():ggplot2中没有theme_blank(),改用theme_bw()或theme_minimal()等标准主题。

内容的提问来源于stack exchange,提问作者BRB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 23:45:34