如何在ggplot美学映射中使用数据集整体计数替代nrow()?
通用化ggplot子组密度总面积归一化方法
之前我实现了一种让堆叠子组密度图总面积为1的方法,代码如下:
library(tidyverse) diamonds %>% ggplot() + geom_density(aes(x = price, y = after_stat(count)/nrow(diamonds), fill = cut), position = "stack") + ylab("Density")
但该写法存在不足:nrow(diamonds)是硬编码的数据集名称,切换到其他数据集时必须手动修改公式。由于ggplot的after_stat()默认仅能访问当前分组的统计量(如count为分组内计数),没有内置的全局总计数变量,以下是几种无需硬编码数据集名的通用解决方案:
方案1:提前提取全局总计数(最简洁)
在绘图前从数据源中提取总观测数,存入变量后在aes中引用:
library(tidyverse) # 替换为任意目标数据集 target_data <- diamonds total_obs <- nrow(target_data) target_data %>% ggplot() + geom_density(aes(x = price, y = after_stat(count)/total_obs, fill = cut), position = "stack") + ylab("Density")
方案2:通过隐藏统计层计算全局总计数
利用stat_summary添加一个不显示的图层,在绘图流程中计算全局总计数,后续在geom_density中引用该值:
library(tidyverse) diamonds %>% ggplot(aes(x = price, fill = cut)) + # 隐藏图层,仅计算全局总计数 stat_summary( aes(y = after_stat(count), global_n = after_stat(max(n()))), geom = "blank", fun = identity, na.rm = TRUE ) + geom_density(aes(y = after_stat(count)/after_stat(global_n)), position = "stack") + ylab("Density")
方案3:自定义统计变换(适合重复使用)
如果需要频繁使用该功能,可以自定义一个统计变换,自动计算全局总计数并归一化:
library(tidyverse) # 自定义统计量 stat_density_global <- function(mapping = NULL, data = NULL, geom = "density", position = "stack", ..., na.rm = FALSE, show.legend = NA, inherit.aes = TRUE) { layer( data = data, mapping = mapping, stat = StatDensityGlobal, geom = geom, position = position, show.legend = show.legend, inherit.aes = inherit.aes, params = list(na.rm = na.rm, ...) ) } StatDensityGlobal <- ggproto("StatDensityGlobal", StatDensity, compute_group = function(data, scales, ...) { # 获取全局总观测数 global_n <- nrow(data) # 调用原密度统计逻辑并归一化 res <- super$compute_group(data, scales, ...) res$y <- res$count / global_n res } ) # 使用自定义统计量绘图 diamonds %>% ggplot() + stat_density_global(aes(x = price, fill = cut), position = "stack") + ylab("Density")
说明
- 方案1是日常使用中最推荐的方式,仅需提前定义一次总计数变量,切换数据集时只需修改
target_data即可。 - 所有方案均无需硬编码数据集名称,适配任意ggplot数据源。
内容的提问来源于stack exchange,提问作者user2554330
相关产品推荐
相关产品推荐

