如何使用dplyr按月份统计各物种的存在个体数量?
用dplyr实现物种月度存在个体数量统计
原始数据
首先定义原始数据框:
data.frame( stringsAsFactors = FALSE, species = c("species1","species1", "species1","species1","species1","species2","species2", "species2","species2","species2"), site = c("A", "B", "C", "D", "E", "A", "B", "C", "D", "E"), first.date = c("01/01/2024","31/01/2024", "02/02/2024","01/03/2024","05/04/2024","20/04/2024", "05/05/2024","15/05/2024","05/06/2024","21/06/2024"), last.date = c("20/02/2024","02/02/2024", "03/02/2024","15/03/2024","03/06/2024","10/05/2024", "20/05/2024","20/05/2024","21/06/2024","21/06/2024") )-> df
数据预览:
| species | site | first.date | last.date |
|---|---|---|---|
| species1 | A | 01/01/2024 | 20/02/2024 |
| species1 | B | 31/01/2024 | 02/02/2024 |
| species1 | C | 02/02/2024 | 03/02/2024 |
| species1 | D | 01/03/2024 | 15/03/2024 |
| species1 | E | 05/04/2024 | 03/06/2024 |
| species2 | A | 20/04/2024 | 10/05/2024 |
| species2 | B | 05/05/2024 | 20/05/2024 |
| species2 | C | 15/05/2024 | 20/05/2024 |
| species2 | D | 05/06/2024 | 21/06/2024 |
| species2 | E | 21/06/2024 | 21/06/2024 |
需求
统计每个物种在每个月份的存在个体数量,预期输出如下:
data.frame( stringsAsFactors = FALSE, species = c("species1","species1", "species1","species1","species1","species1","species2", "species2","species2"), month = c("January","February","March", "April","May","June","April","May","June"), count = c(2L, 3L, 1L, 1L, 1L, 1L, 1L, 3L, 2L) ) -> per_month
预览结果:
| species | month | count |
|---|---|---|
| species1 | January | 2 |
| species1 | February | 3 |
| species1 | March | 1 |
| species1 | April | 1 |
| species1 | May | 1 |
| species1 | June | 1 |
| species2 | April | 1 |
| species2 | May | 3 |
| species2 | June | 2 |
解决方案(dplyr实现)
可以通过dplyr结合lubridate处理日期,生成每个个体覆盖的月份序列后统计数量,代码如下:
library(dplyr) library(lubridate) per_month <- df %>% # 将字符串日期转换为日期格式(日/月/年) mutate( first.date = dmy(first.date), last.date = dmy(last.date) ) %>% # 为每一行生成从first.date到last.date的所有月份序列 rowwise() %>% mutate( month = list(seq(floor_date(first.date, "month"), floor_date(last.date, "month"), by = "month")) ) %>% unnest(month) %>% # 将月份转换为英文全称 mutate(month = month(month, label = TRUE, abbr = FALSE)) %>% # 按物种和月份分组统计数量 group_by(species, month) %>% summarise(count = n(), .groups = "drop") %>% # 按物种和月份顺序排序,匹配预期输出 arrange(species, match(month, month.name))
代码解释
- 日期格式转换:用
lubridate::dmy()将日/月/年格式的字符串转为日期对象; - 生成月份序列:用
floor_date()将日期统一到当月第一天,再用seq()生成覆盖的所有月份,通过rowwise()和unnest()展开每个个体对应的所有月份; - 月份名称转换:用
month()函数将日期格式的月份转为英文全称; - 分组统计:按物种和月份分组,用
n()统计每组的个体数量; - 排序:按物种和月份自然顺序排序,匹配预期输出的格式。
运行上述代码后即可得到符合要求的结果。
内容的提问来源于stack exchange,提问作者luciano
相关产品推荐
相关产品推荐

