如何在R中合并国家数据集并提取每月首个可用日期的死亡数据
R 合并数据集并生成月度首天死亡数字段解决方案
以下是可直接运行的实现代码,适配你的需求:
# 依赖包安装:如果未安装先运行 install.packages(c("tidyverse", "lubridate")) library(tidyverse) library(lubridate) # 用于日期处理,tidyverse部分版本默认不加载,建议单独引入 # -------------------------- 样例数据集构造,你可以替换为自己的数据集 -------------------------- # 人口数据集:对应你说的第一个数据集 df_pop <- tibble( country = c("中国", "美国", "日本"), population = c(1411750000, 333287557, 125700000) ) # 死亡统计数据集:对应你说的第二个数据集 df_death <- tibble( country = rep(c("中国", "美国", "日本"), each = 20), date = seq.Date(as.Date("2023-01-01"), as.Date("2023-06-30"), by = "day") %>% sample(60, replace = T), deaths = sample(10:1000, 60, replace = T) ) # ------------------------------------------------------------------------------------------ # 核心处理逻辑 df_result <- df_death %>% # 标准化日期格式,生成年月分组标识 mutate( date = ymd(date), year_month = floor_date(date, unit = "month") ) %>% # 按国家+年月分组,取每组日期最早的一条死亡记录 group_by(country, year_month) %>% arrange(date) %>% slice_head(n = 1) %>% ungroup() %>% # 长表转宽表,生成单独的月度死亡数字段 pivot_wider( id_cols = country, names_from = year_month, values_from = deaths, # 自定义字段名格式,以下配置生成 deaths_202301 这类格式的字段,可按需调整 names_glue = "deaths_{format(year_month, '%Y%m')}" ) %>% # 和人口数据集按国家关联 left_join(df_pop, by = "country") %>% # 调整字段顺序,人口字段放在最前 select(country, population, everything())
处理完成的输出格式(以2个月份为例)如下:
| country | population | deaths_202301 | deaths_202302 |
|---|---|---|---|
| 中国 | 1411750000 | 123 | 145 |
| 美国 | 333287557 | 456 | 432 |
| 日本 | 125700000 | 78 | 92 |
注意事项
- 如果某个国家某个月无任何死亡记录,对应字段会自动填充为
NA,可后续通过replace_na函数自定义填充值 - 无需提前限定月份范围,代码会自动提取死亡数据集中所有存在的月份生成对应字段
- 若需用base R实现,核心逻辑为用
aggregate取每月最小日期的记录,再用reshape转宽表后合并即可
内容的提问来源于stack exchange,提问作者Posh
相关产品推荐
相关产品推荐

