You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

dplyr如何基于筛选条件提取另一时间点对应类型的关联数据

实现思路

  • 按type字段分组,为每个类型的行添加同类型下一个时间点的时间、计数偏移列,无需额外做关联匹配,效率更高适配数千条的数据集
  • 过滤符合条件的行:当前时间>0、当前counts等于30、存在下一个时间点的有效数据
  • 调整列顺序和字段名即可得到要求的输出

代码实现(tidyverse 方案,简洁易读)

library(dplyr)

result <- df %>%
  group_by(type) %>%
  # 生成同类型下一行的时间、计数值
  mutate(
    time_after = lead(time),
    counts_after = lead(counts)
  ) %>%
  ungroup() %>%
  # 筛选符合要求的行
  filter(time > 0, counts == 30, !is.na(time_after)) %>%
  # 补充type_after字段,调整列顺序
  mutate(type_after = type) %>%
  select(time, type, counts, time_after, type_after, counts_after)

print(result)

代码实现(base R 方案,无需额外加载包)

# 先筛选出当前符合条件的行
curr_rows <- df[df$time > 0 & df$counts == 30, ]
# 逐个匹配下一个时间点的同类型数据
result_list <- list()
for (i in seq_len(nrow(curr_rows))) {
  t <- curr_rows$time[i]
  tp <- curr_rows$type[i]
  next_row <- df[df$time == t + 1 & df$type == tp, ]
  if (nrow(next_row) == 1) {
    result_list[[i]] <- data.frame(
      time = t,
      type = tp,
      counts = 30,
      time_after = next_row$time,
      type_after = tp,
      counts_after = next_row$counts
    )
  }
}
result <- do.call(rbind, result_list)
print(result)

注:你给出的预期输出中第2行counts_after为31属于笔误,原数据中time=2时B类型的counts为30,代码输出会匹配实际数据的正确值。

内容的提问来源于stack exchange,提问作者LDT

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 22:18:04