You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R dplyr实操:基于连续TRUE/FALSE列为聚类组添加分组编号

解决方案

你可以通过状态切换计数的逻辑实现聚类编号的生成,完整dplyr实现代码如下:

library(dplyr)

# 原始数据构造
dat <- data.frame(
  bar = c(letters[1:10]),
  foo = c(1,2,3,5,8,9,11,13,14,15)
)

# 生成in_cluster列+cluster_number列
dat <- dat %>%
  mutate(
    in_cluster = lead(foo) == foo +1 | lag(foo) == foo -1,
    # 识别聚类起始点:当前是TRUE,前一个是FALSE(首行前值默认设为FALSE)
    cluster_start = in_cluster & !lag(in_cluster, default = FALSE),
    # 累加起始点计数得到聚类编号,非聚类行替换为NA
    cluster_number = ifelse(in_cluster, cumsum(cluster_start), NA)
  ) %>%
  # 移除辅助列
  select(-cluster_start)

运行后输出的dat和预期结果完全一致:

bar foo in_cluster cluster_number
1    a   1       TRUE              1
2    b   2       TRUE              1
3    c   3       TRUE              1
4    d   5      FALSE             NA
5    e   8       TRUE              2
6    f   9       TRUE              2
7    g  11      FALSE             NA
8    h  13       TRUE              3
9    i  14       TRUE              3
10   j  15       TRUE              3

逻辑说明

  • 用lag(in_cluster, default = FALSE)取上一行的聚类状态,首行默认赋值为FALSE避免NA值干扰
  • 当当前行in_cluster为TRUE、上一行为FALSE时,说明是新聚类的起始点,标记为TRUE
  • 对所有起始点做累加求和,每遇到一个新起始点计数+1,刚好对应不同聚类的递增编号
  • 最后将in_cluster为FALSE的行的编号替换为NA即可

内容的提问来源于stack exchange,提问作者MKR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 09:24:05