You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr为不同类型分组分配独立编号,非适用组设0

使用dplyr为分组分配独立编号("Not applicable"设为0)

问题描述

给定如下数据:

example_data <- 
  data.frame(value = c(1,3,4,6,7,8,4,6,9,0),
             group = c("Not applicable",
                       "Large group",
                       "Large group",
                       "Not applicable",
                       "Group of 1",
                       "Large group",
                       "Large group",
                       "Large group",
                       "Group of 1",
                       "Not applicable"))

需要实现:

  • 为"Large group"和"Group of 1"的每个连续块分配从1开始的独立递增编号
  • "Not applicable"对应的编号固定为0
  • 其中"Not applicable"可连续多行,"Group of 1"始终单行,"Large group"可任意行数

期望输出:

value          group group_number
1      1 Not applicable            0
2      3    Large group            1
3      4    Large group            1
4      6 Not applicable            0
5      7     Group of 1            2
6      8    Large group            3
7      4    Large group            3
8      6    Large group            3
9      9     Group of 1            4
10     0 Not applicable            0

之前尝试的代码将"Large group"和"Group of 1"视为同一类,导致编号重复:

example_data %>%
  mutate(group_number = with(rle(group != "Not applicable"), 
                      rep(cumsum(values) * values, lengths)))

解决方案

方法1:基于rle()函数实现

通过识别每个连续的分组块,为非"Not applicable"的块分配独立编号:

library(dplyr)

example_data %>%
  mutate(
    group_number = with(rle(group), {
      # 为每个连续组块分配编号:Not applicable为0,其余按出现顺序递增
      block_ids = ifelse(values == "Not applicable", 0, 
                        which(values != "Not applicable"))
      # 将编号按组块长度重复,映射到每一行
      rep(block_ids, lengths)
    })
  )

思路说明

  1. rle(group)返回连续分组的核心信息:values是每个连续块的组名,lengths是每个块的行数
  2. 创建block_ids向量:对"Not applicable"块赋值0,对其他块按出现顺序分配1、2、3...的编号
  3. 用rep(block_ids, lengths)将每个块的编号重复对应行数,得到每行的group_number

方法2:基于dplyr窗口函数实现

无需依赖rle(),用窗口函数标记新组块并累计编号:

library(dplyr)

example_data %>%
  mutate(
    # 标记新的非Not applicable组块的起始行
    new_block = group != "Not applicable" & (row_number() == 1 | group != lag(group)),
    # 累计非Not applicable组块的数量,Not applicable行赋值0
    group_number = cumsum(new_block) * (group != "Not applicable")
  )

思路说明

  1. new_block标记规则:当前行属于非"Not applicable",且要么是第一行,要么和上一行组名不同(即新组块的开始)
  2. cumsum(new_block)累计新组块的数量,再乘以(group != "Not applicable"),将"Not applicable"行的编号置为0

验证结果

两种方法都能得到符合期望的输出:

value          group group_number
1      1 Not applicable            0
2      3    Large group            1
3      4    Large group            1
4      6 Not applicable            0
5      7     Group of 1            2
6      8    Large group            3
7      4    Large group            3
8      6    Large group            3
9      9     Group of 1            4
10     0 Not applicable            0

内容的提问来源于stack exchange,提问作者Polina B

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 03:45:37