如何为循环绘制的各物种小提琴图添加观测数量
问题:为单个物种小提琴图添加观测数量(n值)
数据与初始代码
用户的数据集和初始绘图代码如下:
df = data.frame(c("LALO","LALO", "AMGP", "AMGP", "WRSA", "WRSA", "SNOW", "SNOW"), c(4.3, 2.5, 1.5, 1.2, 0.4, 0.2, 0.1, 0.05)) colnames(df) = c("sp_code","ind_km2") df <- mutate(df, sp_code = factor(sp_code)) # 初始循环绘图代码 sp_code <- levels(unique(df$sp_code)) for (i in seq_along(sp_code)) { species <- sp_code[i] plot2 <- df %>% ggplot(aes(x=sp_code, y=ind_km2, fill=sp_code)) + geom_violin(data = filter(df, sp_code == species), scale = "width", show.legend = FALSE) + stat_summary(data = filter(df, sp_code == species), fun.y=mean, geom="point", shape=23, size=4, fill = "black", show.legend = FALSE) print(plot2) }
用户尝试自定义add_n函数添加n值时,出现错误:geom_text() requires the following missing aesthetics: y。
解决方法
方法一:修复自定义stat_summary函数
原函数存在两个核心问题:引用全局变量导致上下文冲突,以及错误计算观测数(用了全数据集行数而非单个物种行数)。修正后的代码如下:
# 修正后的自定义函数,仅依赖传入的y值计算 add_n <- function(y) { n <- length(y) mean_y <- mean(y) data.frame( y = mean_y * 1.15, # 将n值放在均值的1.15倍位置 label = paste('n =', n) ) } # 循环绘图 sp_code <- levels(unique(df$sp_code)) for (i in seq_along(sp_code)) { species <- sp_code[i] plot2 <- df %>% filter(sp_code == species) %>% # 提前筛选数据,避免重复过滤 ggplot(aes(x=sp_code, y=ind_km2, fill=sp_code)) + geom_violin(scale = "width", show.legend = FALSE) + stat_summary(fun.y=mean, geom="point", shape=23, size=4, fill = "black", show.legend = FALSE) + stat_summary( fun.data = add_n, geom = "text", hjust = 0.5, vjust = 0.9 ) print(plot2) }
方法二:提前统计n值,用annotate添加(更直观)
先统计每个物种的观测数和y轴显示位置,循环中直接用annotate添加文本,逻辑更清晰:
# 提前统计每个物种的n值和推荐的y轴位置(均值的1.15倍) sp_n <- df %>% group_by(sp_code) %>% summarise( n = n(), y_pos = mean(ind_km2) * 1.15 ) # 循环绘图 sp_code <- levels(unique(df$sp_code)) for (i in seq_along(sp_code)) { species <- sp_code[i] current_n <- sp_n %>% filter(sp_code == species) plot2 <- df %>% filter(sp_code == species) %>% ggplot(aes(x=sp_code, y=ind_km2, fill=sp_code)) + geom_violin(scale = "width", show.legend = FALSE) + stat_summary(fun.y=mean, geom="point", shape=23, size=4, fill = "black", show.legend = FALSE) + annotate( "text", x = species, y = current_n$y_pos, label = paste('n =', current_n$n), hjust = 0.5, vjust = 0.9 ) print(plot2) }
关键说明
- 方法一通过让
add_n仅依赖传入的y向量,避免了全局变量的上下文冲突,同时用length(y)准确获取单个物种的观测数。 - 方法二提前统计n值,适合需要灵活自定义y轴显示位置的场景,代码可读性更强。
- 两种方法都提前对数据进行筛选,减少了ggplot内部重复过滤的冗余操作。
内容的提问来源于stack exchange,提问作者Mag
相关产品推荐
相关产品推荐

