You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用ggplot按分组绘制y变量的累计和折线图?

按分组计算累计计数并绘制折线图解决方案

要实现分组随时间的累计计数折线图,核心是先按分组字段分组,再在组内按年份排序后计算累计和,而非直接全局应用cumsum()。以下是具体步骤和代码示例:

1. 数据预处理(dplyr)

假设你的数据集名为df,分组字段为Name.meaning(若需按Name分组,直接替换即可):

情况1:原始数据为单条记录对应一个发布实例

library(dplyr)

df_cum <- df %>%
  # 按目标分组字段分组
  group_by(Name.meaning) %>%
  # 按年份升序排列,保证累计顺序正确
  arrange(Year.published) %>%
  # 计算组内累计计数:每一行计1,逐行累计求和
  mutate(cumulative = cumsum(1)) %>%
  # 取消分组,避免后续ggplot操作出现分组冲突
  ungroup()

情况2:已统计每年每组的数量(含n字段)

如果已经通过count()得到了Year.published、Name.meaning、n的汇总数据:

df_cum <- df %>%
  group_by(Name.meaning) %>%
  arrange(Year.published) %>%
  # 对组内的年度计数n做累计求和
  mutate(cumulative = cumsum(n)) %>%
  ungroup()

可选:补全年份缺失值(避免折线断点)

若部分年份某分组无数据,折线会出现断点,可通过complete()补全并填充0:

df_cum <- df %>%
  group_by(Name.meaning, Year.published) %>%
  summarise(n = n(), .groups = "drop_last") %>%
  # 补全该分组覆盖的所有连续年份,无数据的年份n设为0
  complete(Year.published = full_seq(Year.published, 1), fill = list(n = 0)) %>%
  mutate(cumulative = cumsum(n)) %>%
  ungroup()

2. 绘制分组累计折线图(ggplot2)

用预处理后的df_cum绘图:

library(ggplot2)

ggplot(df_cum, aes(x = Year.published, y = cumulative, color = Name.meaning)) +
  geom_line(linewidth = 1) + # 可根据需求调整线条粗细
  labs(
    x = "发布年份",
    y = "累计计数",
    color = "名称含义",
    title = "各名称含义分组的累计发布数量趋势"
  ) +
  theme_minimal()

关键注意点

  • 必须先通过group_by()指定分组字段,再调用cumsum(),否则会计算全局累计值。
  • 一定要用arrange(Year.published)确保年份按升序排列,否则累计顺序会出错。
  • 若需对比不同分组的累计趋势,将color映射到分组字段即可自动生成分组折线。

内容的提问来源于stack exchange,提问作者harry1027

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 10:54:20