You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于UserID与30天内日期条件合并两个R语言DataFrame的实现方法

R语言DataFrame按匹配规则合并方案

步骤1:转换日期字段格式

首先需要将两个表中存储为字符串的日期字段转换为可计算的日期时间类型:

# 加载依赖包
library(dplyr)
library(lubridate)

# 转换df1的DateTime为日期时间格式
df1 <- df1 %>%
  mutate(DateTime = mdy_hm(DateTime))

# 转换df2的EncounterDate为日期格式
df2 <- df2 %>%
  mutate(EncounterDate = mdy(EncounterDate))

步骤2:按规则合并

方法1:小数据量场景(先关联后过滤)

先按UserID关联两个表,再过滤出日期间隔不超过30天的记录即可:

merged_df <- df1 %>%
  left_join(df2, by = "UserID", suffix = c("_df1", "_df2")) %>%
  mutate(date_gap = abs(as.Date(DateTime) - EncounterDate)) %>%
  filter(date_gap <= days(30)) %>%
  select(-date_gap) # 可选删除临时计算的间隔列

运行后你提到的两条UserID=1的John Smith记录里,只有2021-01-02的那条会匹配成功,2021-12-11的记录因和df2中EncounterDate间隔超过30天会被过滤,完全符合需求。
如果需要保留df1中未匹配到的记录,可将过滤条件改为filter(is.na(date_gap) | date_gap <= days(30))。

方法2:大数据量场景(非等值连接,效率更高)

如果数据量较大,全量关联再过滤会占用过多内存,可直接用非等值连接按条件匹配:

library(powerjoin)
merged_df <- power_left_join(
  df1, df2,
  by = c(
    "UserID",
    ~ abs(as.Date(.x$DateTime) - .y$EncounterDate) <= 30
  ),
  suffix = c("_df1", "_df2")
)

内容的提问来源于stack exchange,提问作者Ash S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 09:45:08