You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中对时间连续的Value值进行正确分组?

R中时间连续条目分组问题解决

问题说明

需要对时间上连续的Value条目进行分组,但现有代码仅能标记连续条目(用"yes"标识),无法区分不同的连续组——不同连续组会被分配同一个分组ID,导致分组失效。

当前代码及输出

当前使用的代码:

df %>%
  mutate(contiguous = ifelse(Endtime_ms == lead(Starttime_ms)|Starttime_ms == lag(Endtime_ms), "yes", "no"),
         grp = consecutive_id(contiguous)
  ) 

输出结果:

# A tibble: 20 × 5
   Value                            Starttime_ms Endtime_ms contiguous   grp
   <chr>                                   <dbl>      <dbl> <chr>      <int>
 1 "on this"                                 210        780 NA             1
 2 "okay"                                   3403       3728 no             2
 3 "cool thanks everyone um"                4221       5880 no             2
 4 "so yes in"                              5910       6900 yes            3 # 一组
 5 "terms of our"                           6900       8370 yes            3 # 一组
 6 "partnership"                            8370       8970 yes            3 # 一组
 7 "projects"                               8970       9480 yes            3 # 一组
 8 "what have we"                           9510      10080 yes            3 # 另一组
 9 "got on the"                            10080      11293 yes            3 # 另一组
10 "horizon? "                             11293      11960 yes            3 # 另一组
11 "let's have a look so the"              11980      13740 no             4
12 "LGBTQ plus"                            13813      16110 no             4
13 "city labs"                             16260      17070 yes            5
14 "have now"                              17070      17910 yes            5
15 "been um"                               17940      19320 no             6
16 "agreed in"                             19350      20190 yes            7
17 "terms of the"                          20190      20760 yes            7
18 "date so"                               20760      21330 yes            7
19 "we're looking at the fifteenth"        21330      22530 yes            7
20 "sixteenth"                             22860      23490 NA             8

期望输出

Value                            Starttime_ms Endtime_ms contiguous   grp
   <chr>                                   <dbl>      <dbl> <chr>      <int>
 1 "on this"                                 210        780 NA             1
 2 "okay"                                   3403       3728 no             2
 3 "cool thanks everyone um"                4221       5880 no             2
 4 "so yes in"                              5910       6900 yes            3 
 5 "terms of our"                           6900       8370 yes            3 
 6 "partnership"                            8370       8970 yes            3 
 7 "projects"                               8970       9480 yes            3 
 8 "what have we"                           9510      10080 yes            4
 9 "got on the"                            10080      11293 yes            4
10 "horizon? "                             11293      11960 yes            4
11 "let's have a look so the"              11980      13740 no             4
12 "LGBTQ plus"                            13813      16110 no             5
13 "city labs"                             16260      17070 yes            6
14 "have now"                              17070      17910 yes            6
15 "been um"                               17940      19320 no             7
16 "agreed in"                             19350      20190 yes            8
17 "terms of the"                          20190      20760 yes            8
18 "date so"                               20760      21330 yes            8
19 "we're looking at the fifteenth"        21330      22530 yes            8
20 "sixteenth"                             22860      23490 NA             9

测试数据

df <- structure(list(Value = c("on this", "okay", "cool thanks everyone um", 
                               "so yes in", "terms of our", "partnership", "projects", "what have we", 
                               "got on the", "horizon? ", "let's have a look so the", "LGBTQ plus", 
                               "city labs", "have now", "been um", "agreed in", "terms of the", 
                               "date so", "we're looking at the fifteenth", "sixteenth"), Starttime_ms = c(210, 
                                                                                                           3403, 4221, 5910, 6900, 8370, 8970, 9510, 10080, 11293, 11980, 
                                                                                                           13813, 16260, 17070, 17940, 19350, 20190, 20760, 21330, 22860
                               ), Endtime_ms = c(780, 3728, 5880, 6900, 8370, 8970, 9480, 10080, 
                                                 11293, 11960, 13740, 16110, 17070, 17910, 19320, 20190, 20760, 
                                                 21330, 22530, 23490)), row.names = c(NA, -20L), class = c("tbl_df", 
                                                                                                           "tbl", "data.frame"))

解决方法

之前的consecutive_id(contiguous)仅根据contiguous的字符值变化生成分组,不同连续组都标记为"yes",因此无法拆分。正确做法是直接判断当前条目与上一条目是否连续,再通过累积求和生成独立分组ID。

修正后代码

library(dplyr)

df %>%
  mutate(
    # 标记当前行是否与上一行时间连续
    is_contiguous = Starttime_ms == lag(Endtime_ms),
    # 生成分组ID:每次不连续时(含第一行)分组序号加1
    grp = cumsum(if_else(is.na(is_contiguous), TRUE, !is_contiguous)),
    # 保留原有的contiguous标记逻辑
    contiguous = case_when(
      is.na(is_contiguous) ~ NA_character_,
      is_contiguous | Endtime_ms == lead(Starttime_ms) ~ "yes",
      TRUE ~ "no"
    )
  ) %>%
  select(Value, Starttime_ms, Endtime_ms, contiguous, grp)

代码解释

  1. is_contiguous:判断当前条目Starttime_ms是否等于上一条目Endtime_ms,确认时间连续性。
  2. grp:用cumsum累积求和,第一行is_contiguous为NA,转为TRUE触发首次计数;后续只要当前条目与上一条不连续,就累加1,确保每个连续组有唯一ID。
  3. contiguous:保留原标记逻辑,同时处理NA值情况。

运行后即可得到符合预期的分组结果。

内容的提问来源于stack exchange,提问作者Chris Ruehlemann

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 09:37:03