You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr为字符向量中重复值序列分配索引

问题:使用dplyr为特定序列分配递增索引

数据

df0 <-
  data.frame(id = 1:13,
             type = c("not_applicable",
                      "new",
                      "same",
                      "same",
                      "not_applicable",
                      "new",
                      "new",
                      "new",
                      "same",
                      "same",
                      "not_applicable",
                      "new",
                      "same"))

需求说明

需要实现以下索引分配规则:

  • 所有not_applicable类型的行,索引设为0
  • 每个new类型的行,获取递增的连续索引(从1开始,依次为1、2、3...)
  • 每个same类型的行,继承前一个new行的索引
  • 序列始终以new开头,若长度大于1则后续跟same

期望输出

data.frame(id = 1:13,
           type = c("not_applicable",
                    "new",
                    "same",
                    "same",
                    "not_applicable",
                    "new",
                    "new",
                    "new",
                    "same",
                    "same",
                    "not_applicable",
                    "new",
                    "same"),
           index = c(0, 1, 1, 1, 0, 2, 3, 4, 4, 4, 0, 5, 5))

尝试的代码及问题

尝试了以下代码:

df0 %>% 
  mutate(index = with(rle(type != "not_applicable"),
                      rep(cumsum(values) * values, lengths)))

得到的输出不符合预期,连续的new行被分配了同一个索引,而非递增的索引:

id           type index
1   1 not_applicable     0
2   2            new     1
3   3           same     1
4   4           same     1
5   5 not_applicable     0
6   6            new     2
7   7            new     2
8   8            new     2
9   9           same     2
10 10           same     2
11 11 not_applicable     0
12 12            new     3
13 13           same     3

解决方案

可以结合dplyr和tidyr的fill函数实现需求,代码如下:

library(dplyr)
library(tidyr)

df0 %>%
  # 为每个new分配递增索引,其他类型先设为NA
  mutate(index = ifelse(type == "new", cumsum(type == "new"), NA)) %>%
  # 向下填充NA,让same继承前一个new的索引
  fill(index, .direction = "down") %>%
  # 将not_applicable对应的index替换为0
  mutate(index = replace(index, type == "not_applicable", 0))

执行后将得到符合预期的输出,其中:

  1. cumsum(type == "new")会为每个new行生成连续递增的数值
  2. fill(index, .direction = "down")将same行的NA值替换为前一个非NA的索引值
  3. 最后用replace把not_applicable行的索引设为0

内容的提问来源于stack exchange,提问作者Polina B

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 12:50:11