在R中按分组基于最小timestamp创建首个A事件行标识新变量
核心判断逻辑很简单,first_A_row仅需同时满足两个条件:
- 当前行
event取值为A - 当前行的
timestamp等于所属participant_id分组下所有A事件对应的最小时间戳
该逻辑无需提前对数据集按时间排序,完全适配你提到的「数据不一定按升序排列」的要求。
方法1:使用dplyr(tidyverse套件)实现
library(dplyr) # 构造示例数据集 df <- data.frame( participant_id = c("ps1", "ps1", "ps1", "ps1", "ps2", "ps2", "ps3", "ps3", "ps3", "ps3"), timestamp = c(0.01, 0.02, 0.03, 0.04, 0.01, 0.02, 0.01, 0.02, 0.03, 0.04), event = c("A", "A", "A", "B", "B", "A", "A", "A", "B", "A") ) # 新增first_A_row列 df <- df %>% group_by(participant_id) %>% mutate(first_A_row = event == "A" & timestamp == min(timestamp[event == "A"], na.rm = TRUE)) %>% ungroup()
方法2:使用基础R实现
无需加载第三方包,直接用内置ave函数实现:
df$first_A_row <- with(df, event == "A" & timestamp == ave( ifelse(event == "A", timestamp, NA), participant_id, FUN = function(x) min(x, na.rm = TRUE) ))
执行后得到的结果和你期望的输出完全一致。
内容的提问来源于stack exchange,提问作者elmolibra
相关产品推荐
相关产品推荐

