如何用dplyr统计同id同天值相同的连续时隙序列长度
dplyr实现连续序列长度统计方案
核心思路
满足同一id、同一day、value值相等的前提下,统计连续递增timeslot序列的长度,实现逻辑如下:
- 按
id、day、value三个维度分组,锁定符合基础条件的统计范围 - 分组内按
timeslot升序排序,确保时间序列顺序正确 - 生成连续段标记:判断当前
timeslot是否等于上一条记录的timeslot+1,不相等则标记为新连续段的起点,通过累积求和生成唯一的连续段ID - 按连续段ID二次分组,统计每组的行数即为该连续序列的长度
t
完整实现代码
# 加载依赖包 library(dplyr) # 读入样本数据 df <- structure(list(id = c(1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1), variable = c("ha1_001", "ha1_002", "ha1_003", "ha1_004", "ha1_125", "ha1_126", "ha1_127", "ha1_128", "ha1_009", "ha1_010", "ha1_011", "ha1_012", "ha1_013"), value = c(110, 110, 110, 110, 110, 110, 110, 110, 110, 110, 110, 110, 110), day = c(1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1), timeslot = c(1, 2, 3, 4, 125, 126, 127, 128, 129, 130, 131, 132, 133), n = c(7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7)), class = c("spec_tbl_df", "tbl_df", "tbl", "data.frame"), row.names = c(NA, -13L), spec = structure(list(cols = list(id = structure(list(), class = c("collector_double", "collector")), variable = structure(list(), class = c("collector_character", "collector")), value = structure(list(), class = c("collector_double", "collector")), day = structure(list(), class = c("collector_double", "collector")), timeslot = structure(list(), class = c("collector_double", "collector")), n = structure(list(), class = c("collector_double", "collector"))), default = structure(list(), class = c("collector_guess", "collector")), skip = 1L), class = "col_spec")) # 计算连续序列长度t df_result <- df %>% # 按id、day、value分组 group_by(id, day, value) %>% # 按timeslot升序排序 arrange(timeslot, .by_group = TRUE) %>% # 生成连续段标记 mutate(consec_group = cumsum(timeslot != lag(timeslot, default = first(timeslot) - 1) + 1)) %>% # 按连续段二次分组 group_by(consec_group, .add = TRUE) %>% # 统计连续段长度 mutate(t = n()) %>% # 移除辅助列 ungroup() %>% select(-consec_group)
结果验证
样本数据运行后会得到3段连续序列:
- timeslot 1-4:t值为4
- timeslot 125-128:t值为4
- timeslot 129-133:t值为5
完全符合需求描述的统计规则。
内容的提问来源于stack exchange,提问作者Rfanatic
相关产品推荐
相关产品推荐

