You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何确定Type A与Type B时间区间的首次精确重叠时刻?

解决Type A与Type B记录的首次重叠时刻问题

方法思路

要找出每条Type A记录与Type B记录的首次精确重叠时刻,核心逻辑是:

  • 匹配所有A与B的组合,筛选出存在时间重叠的记录
  • 计算每组重叠记录的起始时刻(即两个时间段的最晚开始时间,这就是首次重叠的精确时间点)
  • 对每条A记录,保留最早的那个重叠起始时刻

R代码实现(基于dplyr/tidyr)

首先加载所需包并导入数据:

library(dplyr)
library(tidyr)

# 导入数据集
df <- structure(list(type = c("a", "a", "a", "a", "a", "a", "a", "a",
"a", "a", "a", "a", "a", "a", "b", "b", "b", "b", "b", "b", "b",
"b", "b", "b", "b", "b", "b"), starttime = c(470, 858, 1330,
942, 1084, 1320, 1374, 1817, 1394, 1469, 1561, 1796, 1880, 1882,
508, 852, 1203, 1244, 1579, 1865, 2287, 3163, 3784, 4266, 4565,
4936, 5448), endtime = c(485, 873, 1345, 957, 1099, 1335, 1389,
1832, 1409, 1484, 1576, 1811, 1895, 1897, 536, 919, 1216, 1285,
1598, 1892, 2355, 3229, 3817, 4303, 4626, 4976, 5497)), row.names = c(NA,
-27L), class = c("tbl_df", "tbl", "data.frame"))

基础方案(仅保留有重叠的A记录)

result <- df %>%
  # 拆分A和B数据,交叉连接所有组合
  filter(type == "a") %>%
  cross_join(filter(df, type == "b"), suffix = c("_a", "_b")) %>%
  # 筛选时间重叠的记录:A的时间段与B的时间段有交集
  filter(starttime_a <= endtime_b & starttime_b <= endtime_a) %>%
  # 计算首次重叠时刻:取两个时间段的最晚开始时间
  mutate(overlap_start = pmax(starttime_a, starttime_b)) %>%
  # 按每条A记录分组,保留最早的重叠时刻
  group_by(starttime_a, endtime_a) %>%
  slice_min(overlap_start, n = 1) %>%
  ungroup() %>%
  # 整理输出列
  select(
    type_a = type_a,
    a_start = starttime_a,
    a_end = endtime_a,
    first_overlap_time = overlap_start,
    type_b = type_b,
    b_start = starttime_b,
    b_end = endtime_b
  )

print(result)

进阶方案(保留所有A记录,无重叠时标记NA)

如果需要保留所有Type A记录,即使没有对应的B重叠记录,可通过右连接补充:

result_with_na <- df %>%
  filter(type == "a") %>%
  mutate(a_id = row_number()) %>%
  cross_join(filter(df, type == "b"), suffix = c("_a", "_b")) %>%
  filter(starttime_a <= endtime_b & starttime_b <= endtime_a) %>%
  mutate(overlap_start = pmax(starttime_a, starttime_b)) %>%
  group_by(a_id) %>%
  slice_min(overlap_start, n = 1) %>%
  ungroup() %>%
  # 右连接原始A数据,保留所有A记录
  right_join(filter(df, type == "a") %>% mutate(a_id = row_number()), by = "a_id") %>%
  # 整理列并填充NA
  select(
    type_a = type.x,
    a_start = starttime.x,
    a_end = endtime.x,
    first_overlap_time = overlap_start,
    type_b,
    b_start = starttime_b,
    b_end = endtime_b
  ) %>%
  mutate(across(c(first_overlap_time, type_b, b_start, b_end), ~replace_na(., NA)))

print(result_with_na)

高效方案(基于data.table,适合大数据集)

如果数据集规模较大,使用data.table可显著提升处理效率:

library(data.table)

setDT(df)
dt_a <- df[type == "a"]
dt_b <- df[type == "b"]

result_dt <- dt_a[dt_b, on = .(starttime <= endtime, endtime >= starttime), allow.cartesian = TRUE] %>%
  .[, overlap_start := pmax(starttime, i.starttime)] %>%
  .[, .SD[which.min(overlap_start)], by = .(type, starttime, endtime)] %>%
  setnames(
    old = c("type", "starttime", "endtime", "i.type", "i.starttime", "i.endtime", "overlap_start"),
    new = c("type_a", "a_start", "a_end", "type_b", "b_start", "b_end", "first_overlap_time")
  )

print(result_dt)

关键说明

  • 重叠判断条件:starttime_a <= endtime_b & starttime_b <= endtime_a是判断两个时间段存在交集的标准逻辑
  • 首次重叠时刻:取两个时间段的最晚开始时间(pmax(starttime_a, starttime_b)),这是两段时间开始重叠的精确时刻
  • 若多条B记录对应同一个最早重叠时刻,slice_min会保留第一条,可根据需求调整为保留全部或其他规则

内容的提问来源于stack exchange,提问作者Peter Thome

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 14:04:55