如何使用dplyr统计数据框中A首次连续出现的序列长度
R语言dplyr实现统计列中首段连续A的序列长度
实现方案(适用于dplyr 1.1.0及以上版本)
核心使用dplyr内置的consecutive_id函数识别连续值分组,筛选首个A序列后统计长度:
library(dplyr) # 示例数据框 df <- structure(list(col1 = c("A", "A", "A", "B", "A", "A", "C", "A" )), class = "data.frame", row.names = c(NA, -8L)) df %>% # 为连续相同的col1值分配唯一分组ID mutate(seq_group = consecutive_id(col1)) %>% # 筛选所有值为A的行 filter(col1 == "A") %>% # 取最早出现的A对应的分组 filter(seq_group == min(seq_group)) %>% # 统计该分组的行数即为目标长度 summarise(first_A_seq_length = n())
运行后输出结果:
first_A_seq_length 1 3
低版本dplyr兼容方案
如果dplyr版本低于1.1.0没有consecutive_id,可以用cumsum+lag组合实现相同的连续分组逻辑:
df %>% mutate(seq_group = cumsum(col1 != lag(col1, default = first(col1)))) %>% filter(col1 == "A") %>% filter(seq_group == min(seq_group)) %>% summarise(first_A_seq_length = n())
内容的提问来源于stack exchange,提问作者TarJae
相关产品推荐
相关产品推荐

