基于30天日期范围匹配同RespID行计算DataFrame新列
实现方案
首先将原始数据中的年、月、日字段合并为标准日期格式,再按RespID分组判断每个Col4=A的行是否满足条件即可,全程返回结果均为data.frame结构。
完整可运行代码
library(dplyr) library(lubridate) # 构造示例数据 df <- data.frame( ID = 1:5, Col1 = c("blue", "orange", "red", "yellow", "green"), RespID = c("729Ad", "295gS", "729Ad", "592Jd", "937sa"), Col3 = c(3.2, 6.5, 8.4, 2.9, 3.5), Col4 = c("A", "A", "B", "A", "B"), Year = rep(2021,5), Month = c("April", "April", "April", "March", "May"), Day = c(2,1,20,12,13) ) # 核心计算逻辑 df_result <- df %>% # 拼接转换为可计算的标准日期格式 mutate(full_date = ymd(paste(Year, Month, Day))) %>% # 按RespID分组,同ID的行统一判断 group_by(RespID) %>% mutate(Col5 = case_when( # 所有Col4为B的行统一赋值0 Col4 == "B" ~ 0, # Col4为A的行,判断同组是否存在日期差30天内的B行 Col4 == "A" ~ as.integer(any(Col4 == "B" & abs(difftime(full_date[Col4 == "B"], full_date, units = "days")) <= 30)) )) %>% ungroup() %>% # 移除中间生成的日期字段和原年月日字段,匹配预期输出结构 select(-full_date, -Year, -Month, -Day) # 后续需要仅保留A行时直接执行以下代码即可 df_only_A <- filter(df_result, Col4 == "A")
输出结果验证
生成的df_result完全匹配预期结构:
| ID | Col1 | RespID | Col3 | Col4 | Col5 |
|---|---|---|---|---|---|
| 1 | blue | 729Ad | 3.2 | A | 1 |
| 2 | orange | 295gS | 6.5 | A | 0 |
| 3 | red | 729Ad | 8.4 | B | 0 |
| 4 | yellow | 592Jd | 2.9 | A | 0 |
| 5 | green | 937sa | 3.5 | B | 0 |
内容的提问来源于stack exchange,提问作者piper180
相关产品推荐
相关产品推荐

