如何在R中对时间连续的Value值进行正确分组?
R中时间连续条目分组问题解决
问题说明
需要对时间上连续的Value条目进行分组,但现有代码仅能标记连续条目(用"yes"标识),无法区分不同的连续组——不同连续组会被分配同一个分组ID,导致分组失效。
当前代码及输出
当前使用的代码:
df %>% mutate(contiguous = ifelse(Endtime_ms == lead(Starttime_ms)|Starttime_ms == lag(Endtime_ms), "yes", "no"), grp = consecutive_id(contiguous) )
输出结果:
# A tibble: 20 × 5 Value Starttime_ms Endtime_ms contiguous grp <chr> <dbl> <dbl> <chr> <int> 1 "on this" 210 780 NA 1 2 "okay" 3403 3728 no 2 3 "cool thanks everyone um" 4221 5880 no 2 4 "so yes in" 5910 6900 yes 3 # 一组 5 "terms of our" 6900 8370 yes 3 # 一组 6 "partnership" 8370 8970 yes 3 # 一组 7 "projects" 8970 9480 yes 3 # 一组 8 "what have we" 9510 10080 yes 3 # 另一组 9 "got on the" 10080 11293 yes 3 # 另一组 10 "horizon? " 11293 11960 yes 3 # 另一组 11 "let's have a look so the" 11980 13740 no 4 12 "LGBTQ plus" 13813 16110 no 4 13 "city labs" 16260 17070 yes 5 14 "have now" 17070 17910 yes 5 15 "been um" 17940 19320 no 6 16 "agreed in" 19350 20190 yes 7 17 "terms of the" 20190 20760 yes 7 18 "date so" 20760 21330 yes 7 19 "we're looking at the fifteenth" 21330 22530 yes 7 20 "sixteenth" 22860 23490 NA 8
期望输出
Value Starttime_ms Endtime_ms contiguous grp <chr> <dbl> <dbl> <chr> <int> 1 "on this" 210 780 NA 1 2 "okay" 3403 3728 no 2 3 "cool thanks everyone um" 4221 5880 no 2 4 "so yes in" 5910 6900 yes 3 5 "terms of our" 6900 8370 yes 3 6 "partnership" 8370 8970 yes 3 7 "projects" 8970 9480 yes 3 8 "what have we" 9510 10080 yes 4 9 "got on the" 10080 11293 yes 4 10 "horizon? " 11293 11960 yes 4 11 "let's have a look so the" 11980 13740 no 4 12 "LGBTQ plus" 13813 16110 no 5 13 "city labs" 16260 17070 yes 6 14 "have now" 17070 17910 yes 6 15 "been um" 17940 19320 no 7 16 "agreed in" 19350 20190 yes 8 17 "terms of the" 20190 20760 yes 8 18 "date so" 20760 21330 yes 8 19 "we're looking at the fifteenth" 21330 22530 yes 8 20 "sixteenth" 22860 23490 NA 9
测试数据
df <- structure(list(Value = c("on this", "okay", "cool thanks everyone um", "so yes in", "terms of our", "partnership", "projects", "what have we", "got on the", "horizon? ", "let's have a look so the", "LGBTQ plus", "city labs", "have now", "been um", "agreed in", "terms of the", "date so", "we're looking at the fifteenth", "sixteenth"), Starttime_ms = c(210, 3403, 4221, 5910, 6900, 8370, 8970, 9510, 10080, 11293, 11980, 13813, 16260, 17070, 17940, 19350, 20190, 20760, 21330, 22860 ), Endtime_ms = c(780, 3728, 5880, 6900, 8370, 8970, 9480, 10080, 11293, 11960, 13740, 16110, 17070, 17910, 19320, 20190, 20760, 21330, 22530, 23490)), row.names = c(NA, -20L), class = c("tbl_df", "tbl", "data.frame"))
解决方法
之前的consecutive_id(contiguous)仅根据contiguous的字符值变化生成分组,不同连续组都标记为"yes",因此无法拆分。正确做法是直接判断当前条目与上一条目是否连续,再通过累积求和生成独立分组ID。
修正后代码
library(dplyr) df %>% mutate( # 标记当前行是否与上一行时间连续 is_contiguous = Starttime_ms == lag(Endtime_ms), # 生成分组ID:每次不连续时(含第一行)分组序号加1 grp = cumsum(if_else(is.na(is_contiguous), TRUE, !is_contiguous)), # 保留原有的contiguous标记逻辑 contiguous = case_when( is.na(is_contiguous) ~ NA_character_, is_contiguous | Endtime_ms == lead(Starttime_ms) ~ "yes", TRUE ~ "no" ) ) %>% select(Value, Starttime_ms, Endtime_ms, contiguous, grp)
代码解释
is_contiguous:判断当前条目Starttime_ms是否等于上一条目Endtime_ms,确认时间连续性。grp:用cumsum累积求和,第一行is_contiguous为NA,转为TRUE触发首次计数;后续只要当前条目与上一条不连续,就累加1,确保每个连续组有唯一ID。contiguous:保留原标记逻辑,同时处理NA值情况。
运行后即可得到符合预期的分组结果。
内容的提问来源于stack exchange,提问作者Chris Ruehlemann
相关产品推荐
相关产品推荐

