You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按ID分组计算数据框中事件上次发生的间隔轮次

问题

我有如下数据框:

set.seed(1)
minimalsample <- data.frame(ID=rep(1:5, each= 20), round_number= rep(1:20, 5), event_occurred=rbinom(100, size=1, prob=0.2))

部分数据如下:

IDround_numberevent_occurred
110
120
130
141
150
161
171
180
190
1100
1110

我希望新增一列lagEventOccurred,用于标识距离上次事件发生过去了多少轮次,规则如下:

  • 同一轮次发生的事件不计入,仅统计之前的轮次
  • 需按ID分组处理
  • 若之前从未发生事件,值为Inf

目标数据示例如下:

IDround_numberevent_occurredlagEventOccurred
110Inf
120Inf
130Inf
141Inf
1501
1612
1711
1801
1902
11003
11104

我尝试了以下方法,但似乎需要循环执行多次lag()才能得到结果:

minimalsample %>% dplyr::mutate(firststep = ifelse(event_occurred==1, round_number, Inf) )  %>%
  dplyr::mutate(secondstep = ifelse((lag(firststep) < round_number), lag(firststep), firststep )) %>%
  dplyr::mutate(thirdstep = ifelse((lag(firststep) < round_number), round_number- lag(firststep), round_number-firststep ))

请问如何高效实现计算满足条件的上次事件的间隔轮次?


高效实现方案

可以通过向量化操作避免循环,利用dplyr结合tidyr或zoo包实现,以下是两种可行方案:

方案1:使用dplyr + tidyr

library(dplyr)
library(tidyr)

result <- minimalsample %>%
  group_by(ID) %>%
  # 标记事件发生的轮次,非事件行设为NA
  mutate(event_round = ifelse(event_occurred == 1, round_number, NA)) %>%
  # 向下填充最近的事件轮次,让每个行都能获取到截至当前的最近事件位置
  fill(event_round, .direction = "down") %>%
  # 取上一个位置的最近事件轮次(即当前行之前的最近事件)
  mutate(last_event = lag(event_round)) %>%
  # 计算间隔,无前置事件时设为Inf
  mutate(lagEventOccurred = ifelse(is.na(last_event), Inf, round_number - last_event)) %>%
  # 移除中间辅助列
  select(-event_round, -last_event) %>%
  ungroup()

方案2:使用dplyr + zoo

如果习惯用zoo包的na.locf函数(Last Observation Carried Forward),可以用以下写法:

library(dplyr)
library(zoo)

result <- minimalsample %>%
  group_by(ID) %>%
  mutate(
    event_round = ifelse(event_occurred == 1, round_number, NA),
    # 向前填充最近的事件轮次
    last_event = na.locf(event_round, na.rm = FALSE),
    # 取当前行之前的最近事件轮次
    last_event = lag(last_event),
    # 计算间隔并处理无前置事件的情况
    lagEventOccurred = ifelse(is.na(last_event), Inf, round_number - last_event)
  ) %>%
  select(-event_round, -last_event) %>%
  ungroup()

逻辑说明

  1. 分组处理:按ID分组,确保每个用户的计算独立
  2. 标记事件位置:用event_round记录所有事件发生的轮次,非事件行留空
  3. 填充最近事件:通过fill或na.locf将最近的事件轮次向下传递,让每个行都能获取到截至当前的最近事件
  4. 取前置事件:用lag()将填充后的事件轮次列下移一位,得到当前行之前的最近事件轮次
  5. 计算间隔:用当前轮次减去前置事件轮次,无前置事件时替换为Inf

这两种方法都是向量化操作,无需循环,处理大规模数据时效率远高于循环方案。


内容的提问来源于stack exchange,提问作者Zahra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 13:32:04