You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为ggridges分箱脊线图按物种单独添加百分比标签

解决ggridges分箱脊线图按物种单独计算百分比标签的问题

你当前代码中sum(count)是对全数据集的计数求和,导致百分比是基于整体数据的占比,而非每个物种内部的占比。要实现每个物种的百分比总和为100,需要按物种分组计算组内计数总和,以下是两种可行方案:

方案1:直接在ggplot中调整统计计算

利用ave(count, group, FUN = sum)获取每个物种组内的总计数,替代全局的sum(count),从而计算组内百分比:

library(ggplot2)
library(ggridges)

iris_mod <- rbind(iris, iris[iris$Species == "setosa",])

ggplot(iris_mod, aes(x = Sepal.Length, y = Species, fill = Species)) +
  geom_density_ridges(alpha = 0.6, stat = "binline", binwidth = .5, draw_baseline = FALSE, boundary = 0) +
  geom_text(
    stat = "bin",
    aes(y = group + 0.05, # 微调y轴位置,避免标签与脊线重叠
        label = round(stat(count / ave(count, group, FUN = sum) * 100), 2)),
    vjust = 0, size = 3, color = "black", binwidth = .5, boundary = 0
  )

方案2:提前预处理数据(更直观可控)

先通过分组分箱计算好每个物种的分箱计数和百分比,再传入ggplot绘图,这种方式便于后续调整分箱规则或标签样式:

library(tidyverse)
library(ggridges)

iris_mod <- rbind(iris, iris[iris$Species == "setosa",])

# 预处理数据:按物种分组分箱,计算组内百分比
iris_binned <- iris_mod %>%
  group_by(Species) %>%
  # 按0.5宽度分箱,包含最小值
  mutate(Sepal_bin = cut(Sepal.Length, breaks = seq(floor(min(Sepal.Length)), ceiling(max(Sepal.Length)), by = 0.5), include.lowest = TRUE)) %>%
  count(Sepal_bin) %>%
  # 计算组内百分比并保留两位小数
  mutate(pct = round(n / sum(n) * 100, 2)) %>%
  ungroup() %>%
  # 计算分箱中点,用于x轴标签的位置对齐
  mutate(Sepal_mid = map_dbl(Sepal_bin, ~mean(as.numeric(.x))))

# 绘制脊线图
ggplot(iris_binned, aes(x = Sepal_mid, y = Species, fill = Species)) +
  # 使用stat="identity"直接调用预处理好的计数作为脊线高度
  geom_density_ridges(stat = "identity", aes(height = n), alpha = 0.6, draw_baseline = FALSE) +
  # 添加百分比标签
  geom_text(aes(label = pct), vjust = 0, size = 3, color = "black", nudge_y = 0.05) +
  # 调整x轴刻度与分箱宽度一致
  scale_x_continuous(breaks = seq(4, 8, by = 0.5)) +
  labs(x = "Sepal Length")

内容的提问来源于stack exchange,提问作者user236321

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 02:55:31