基于分组内日期筛选创建新列的R语言技术问询
按ID分组获取指定年份的首个职位解决方案
原始数据
首先是待处理的数据框my_data:
# dput输出 structure(list(id = c(1, 1, 1, 2, 2, 3, 3), begin = c("2017-01-01", "2017-08-01", "2022-05-01", "2017-01-01", "2017-09-01", "2017-01-01", "2017-09-01"), end = c("2017-07-01", "2022-04-01", "2023-06-01", "2017-08-01", "2023-06-01", "2017-08-01", "2023-06-01"), position = c("position_1", "position_2", "position_3", "position_1", "position_2", "position_1", "position_1")), row.names = c(NA, -7L), class = "data.frame") # 打印预览 id begin end position 1 1 2017-01-01 2017-07-01 position_1 2 1 2017-08-01 2022-04-01 position_2 3 1 2022-05-01 2023-06-01 position_3 4 2 2017-01-01 2017-08-01 position_1 5 2 2017-09-01 2023-06-01 position_2 6 3 2017-01-01 2017-08-01 position_1 7 3 2017-09-01 2023-06-01 position_1
需求说明
按id分组,生成如first_position_in_2023、first_position_in_2017这类年份对应列,存储每个ID在目标年份1月1日当天处于的首个职位(即时间区间[begin, end]包含YYYY-01-01的首个position)。
问题分析
你尝试的代码返回了所有行的结果,原因是summarise中直接使用position[begin <= ...]会保留分组内所有满足条件的元素,没有提取首个匹配项,同时也漏了判断end >= 目标日期(确保日期在区间内)。
解决方案
1. 先转换日期类型
首先将begin和end转换为日期格式,保证日期比较的准确性:
library(dplyr) library(lubridate) my_data <- my_data %>% mutate(across(c(begin, end), ymd))
2. 单年份的处理方法
按ID分组后,筛选出包含目标日期的行,提取首个position:
my_data %>% group_by(id) %>% summarise( first_position_in_2023 = position[begin <= ymd("2023-01-01") & end >= ymd("2023-01-01")][1] )
运行结果:
# A tibble: 3 x 2 id first_position_in_2023 <dbl> <chr> 1 1 position_3 2 2 position_2 3 3 position_1
3. 多年份的自动化处理
如果需要批量生成多个年份的列,可以用purrr实现自动化:
library(purrr) # 指定需要处理的年份范围 target_years <- 2017:2023 # 定义生成单年份列的函数 get_year_position <- function(year) { target_date <- ymd(paste0(year, "-01-01")) my_data %>% group_by(id) %>% summarise( !!paste0("first_position_in_", year) := position[begin <= target_date & end >= target_date][1] ) } # 合并所有年份的结果 result <- map_dfc(target_years, get_year_position) %>% select(id, everything()) %>% distinct(id, .keep_all = TRUE)
运行后会得到包含first_position_in_2017到first_position_in_2023所有列的数据框,每个ID对应一行。
内容的提问来源于stack exchange,提问作者fschier
相关产品推荐
相关产品推荐

