You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为循环绘制的各物种小提琴图添加观测数量

问题:为单个物种小提琴图添加观测数量(n值)

数据与初始代码

用户的数据集和初始绘图代码如下:

df = data.frame(c("LALO","LALO", "AMGP", "AMGP", "WRSA", "WRSA", "SNOW", "SNOW"),
                c(4.3, 2.5, 1.5, 1.2, 0.4, 0.2, 0.1, 0.05))
colnames(df) = c("sp_code","ind_km2")

df <- mutate(df, sp_code = factor(sp_code))

# 初始循环绘图代码
sp_code <- levels(unique(df$sp_code))

for (i in seq_along(sp_code)) {
  species <- sp_code[i]
  plot2 <- df %>% 
    ggplot(aes(x=sp_code, y=ind_km2, fill=sp_code)) +
    geom_violin(data = filter(df, sp_code == species), scale = "width", show.legend = FALSE) +
    stat_summary(data = filter(df, sp_code == species), fun.y=mean, geom="point", shape=23, size=4, fill = "black", show.legend = FALSE)
  print(plot2)
} 

用户尝试自定义add_n函数添加n值时,出现错误:geom_text() requires the following missing aesthetics: y。


解决方法

方法一:修复自定义stat_summary函数

原函数存在两个核心问题:引用全局变量导致上下文冲突,以及错误计算观测数(用了全数据集行数而非单个物种行数)。修正后的代码如下:

# 修正后的自定义函数,仅依赖传入的y值计算
add_n <- function(y) {
  n <- length(y)
  mean_y <- mean(y)
  data.frame(
    y = mean_y * 1.15,  # 将n值放在均值的1.15倍位置
    label = paste('n =', n)
  )
}

# 循环绘图
sp_code <- levels(unique(df$sp_code))

for (i in seq_along(sp_code)) {
  species <- sp_code[i]
  plot2 <- df %>%
    filter(sp_code == species) %>%  # 提前筛选数据,避免重复过滤
    ggplot(aes(x=sp_code, y=ind_km2, fill=sp_code)) +
    geom_violin(scale = "width", show.legend = FALSE) +
    stat_summary(fun.y=mean, geom="point", shape=23, size=4, fill = "black", show.legend = FALSE) +
    stat_summary(
      fun.data = add_n,
      geom = "text",
      hjust = 0.5,
      vjust = 0.9
    )
  print(plot2)
}

方法二:提前统计n值,用annotate添加(更直观)

先统计每个物种的观测数和y轴显示位置,循环中直接用annotate添加文本,逻辑更清晰:

# 提前统计每个物种的n值和推荐的y轴位置(均值的1.15倍)
sp_n <- df %>%
  group_by(sp_code) %>%
  summarise(
    n = n(),
    y_pos = mean(ind_km2) * 1.15
  )

# 循环绘图
sp_code <- levels(unique(df$sp_code))

for (i in seq_along(sp_code)) {
  species <- sp_code[i]
  current_n <- sp_n %>% filter(sp_code == species)
  
  plot2 <- df %>%
    filter(sp_code == species) %>%
    ggplot(aes(x=sp_code, y=ind_km2, fill=sp_code)) +
    geom_violin(scale = "width", show.legend = FALSE) +
    stat_summary(fun.y=mean, geom="point", shape=23, size=4, fill = "black", show.legend = FALSE) +
    annotate(
      "text",
      x = species,
      y = current_n$y_pos,
      label = paste('n =', current_n$n),
      hjust = 0.5,
      vjust = 0.9
    )
  print(plot2)
}

关键说明

  • 方法一通过让add_n仅依赖传入的y向量,避免了全局变量的上下文冲突,同时用length(y)准确获取单个物种的观测数。
  • 方法二提前统计n值,适合需要灵活自定义y轴显示位置的场景,代码可读性更强。
  • 两种方法都提前对数据进行筛选,减少了ggplot内部重复过滤的冗余操作。

内容的提问来源于stack exchange,提问作者Mag

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 10:09:52