求助:基于id和增量值标记含中断的TRUE序列
解决R中按ID标记TRUE序列(含中断序列的第一个FALSE)的问题
我来帮你搞定这个序列标记的需求!你的核心要求是:
- 按
id分组处理数据 - 连续的
TRUE构成一个序列,中断该序列的第一个FALSE要归入这个序列 - 那些不关联任何TRUE序列的连续
FALSE,统一标记为0
下面我会给出两种可行的解决方案,一种基于tidyverse工具链(更简洁),另一种基于你尝试的rle()函数(更贴合你最初的思路)。
先准备样本数据
首先把你提供的测试数据加载进来,方便后续验证:
test <- structure(list(id = c(1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3), logical = c(TRUE, TRUE, FALSE, TRUE, TRUE, FALSE, TRUE, TRUE, TRUE, TRUE, FALSE, TRUE, TRUE, TRUE, FALSE, FALSE, FALSE, TRUE, FALSE, TRUE, FALSE, FALSE, FALSE, FALSE, TRUE)), .Names = c("id", "logical"), class = "data.frame", row.names = c(NA, -25L))
方法一:用tidyverse(dplyr + tidyr)实现
这种方法代码简洁,逻辑清晰,适合熟悉tidyverse的用户:
library(dplyr) library(tidyr) test_result <- test %>% group_by(id) %>% mutate( # 第一步:标记哪些行属于某个序列:TRUE本身,或者是TRUE序列后的第一个FALSE in_sequence = logical | (lag(logical, default = FALSE) & !logical), # 第二步:给连续的有效序列块分配临时编号,非序列行标记为0 block_id = cumsum(in_sequence & lag(in_sequence, default = FALSE) == FALSE) * in_sequence ) %>% # 第三步:把临时编号转换成每个id内递增的序列号 group_by(id, block_id) %>% mutate(sequence = ifelse(block_id == 0, 0, cur_group_id())) %>% ungroup() %>% # 保留需要的列 select(id, logical, sequence)
验证结果
比如查看id=3的部分,和你的示例完全匹配:
filter(test_result, id == 3)
输出:
# A tibble: 11 × 3 id logical sequence <dbl> <lgl> <dbl> 1 3 FALSE 0 2 3 FALSE 0 3 3 FALSE 0 4 3 TRUE 1 5 3 FALSE 1 6 3 TRUE 2 7 3 FALSE 2 8 3 FALSE 0 9 3 FALSE 0 10 3 FALSE 0 11 3 TRUE 3
方法二:基于rle()函数实现
既然你已经尝试了rle(),那我们就顺着这个思路,手动处理每个分组的rle结果:
library(dplyr) test_result_rle <- test %>% group_by(id) %>% mutate( sequence = { # 对当前id的logical列做游程编码 r <- rle(logical) seq_num <- integer(length(r$lengths)) current_seq <- 0 # 遍历每个游程块 for(i in seq_along(r$lengths)){ if(r$values[i]){ # 遇到TRUE块,序列号+1 current_seq <- current_seq + 1 seq_num[i] <- current_seq # 检查下一个块是否是单个FALSE(即中断当前序列的那个FALSE) if(i < length(r$lengths) && !r$values[i+1] && r$lengths[i+1] == 1){ seq_num[i+1] <- current_seq i <- i + 1 # 跳过这个FALSE块,避免重复处理 } } else { # 无关的FALSE块标记为0 seq_num[i] <- 0 } } # 把游程编码的结果扩展成和原数据长度一致的向量 rep(seq_num, r$lengths) } ) %>% ungroup()
这个方法的核心是遍历每个id的rle结果,遇到TRUE序列就分配序号,同时判断后续是否有单个FALSE需要归入当前序列,逻辑和你的需求完全对齐。
两种方法都能得到你想要的结果,你可以根据自己的代码习惯选择。如果有其他细节需要调整,随时告诉我!
内容的提问来源于stack exchange,提问作者iskandarblue
相关产品推荐
相关产品推荐

