如何在R中按个体首次记录时间筛选1年内的数据集
解决方案
首先需要将字符型的日期列转换为R可识别的日期格式,之后按个体分组,筛选出每个个体首次记录起1年内的数据。以下提供两种常用实现方式:
方法一:Base R 实现
# 转换日期格式 df_date$Dates <- as.Date(df_date$Dates) # 按Name拆分数据,筛选每个个体的目标数据后合并 result_base <- do.call(rbind, lapply(split(df_date, df_date$Name), function(sub_df) { first_date <- min(sub_df$Dates) # 用+365近似1年,如需精确处理闰年,可替换为seq(first_date, by = "year", length.out = 2)[2] - 1 sub_df[sub_df$Dates >= first_date & sub_df$Dates <= first_date + 365, ] })) # 查看结果 print(result_base)
方法二:Tidyverse 实现
需依赖dplyr和lubridate包处理分组与日期运算:
library(dplyr) library(lubridate) # 转换日期格式并完成筛选 result_tidy <- df_date %>% mutate(Dates = ymd(Dates)) %>% group_by(Name) %>% filter(Dates >= min(Dates) & Dates <= min(Dates) + years(1)) %>% ungroup() # 查看结果 print(result_tidy)
结果验证
对示例数据执行后:
- Jim的保留日期范围为
2010-01-01至2010-12-28(2010-01-01加1年为2011-01-01,2011-01-16超出范围被剔除) - Sue的保留日期范围为
2010-04-01至2011-03-28(2010-04-01加1年为2011-04-01,2012-12-16超出范围被剔除)
内容的提问来源于stack exchange,提问作者John Conor
相关产品推荐
相关产品推荐

