如何用ggplot按分组绘制y变量的累计和折线图?
按分组计算累计计数并绘制折线图解决方案
要实现分组随时间的累计计数折线图,核心是先按分组字段分组,再在组内按年份排序后计算累计和,而非直接全局应用cumsum()。以下是具体步骤和代码示例:
1. 数据预处理(dplyr)
假设你的数据集名为df,分组字段为Name.meaning(若需按Name分组,直接替换即可):
情况1:原始数据为单条记录对应一个发布实例
library(dplyr) df_cum <- df %>% # 按目标分组字段分组 group_by(Name.meaning) %>% # 按年份升序排列,保证累计顺序正确 arrange(Year.published) %>% # 计算组内累计计数:每一行计1,逐行累计求和 mutate(cumulative = cumsum(1)) %>% # 取消分组,避免后续ggplot操作出现分组冲突 ungroup()
情况2:已统计每年每组的数量(含n字段)
如果已经通过count()得到了Year.published、Name.meaning、n的汇总数据:
df_cum <- df %>% group_by(Name.meaning) %>% arrange(Year.published) %>% # 对组内的年度计数n做累计求和 mutate(cumulative = cumsum(n)) %>% ungroup()
可选:补全年份缺失值(避免折线断点)
若部分年份某分组无数据,折线会出现断点,可通过complete()补全并填充0:
df_cum <- df %>% group_by(Name.meaning, Year.published) %>% summarise(n = n(), .groups = "drop_last") %>% # 补全该分组覆盖的所有连续年份,无数据的年份n设为0 complete(Year.published = full_seq(Year.published, 1), fill = list(n = 0)) %>% mutate(cumulative = cumsum(n)) %>% ungroup()
2. 绘制分组累计折线图(ggplot2)
用预处理后的df_cum绘图:
library(ggplot2) ggplot(df_cum, aes(x = Year.published, y = cumulative, color = Name.meaning)) + geom_line(linewidth = 1) + # 可根据需求调整线条粗细 labs( x = "发布年份", y = "累计计数", color = "名称含义", title = "各名称含义分组的累计发布数量趋势" ) + theme_minimal()
关键注意点
- 必须先通过
group_by()指定分组字段,再调用cumsum(),否则会计算全局累计值。 - 一定要用
arrange(Year.published)确保年份按升序排列,否则累计顺序会出错。 - 若需对比不同分组的累计趋势,将
color映射到分组字段即可自动生成分组折线。
内容的提问来源于stack exchange,提问作者harry1027
相关产品推荐
相关产品推荐

