使用dplyr为字符向量中重复值序列分配索引
问题:使用dplyr为特定序列分配递增索引
数据
df0 <- data.frame(id = 1:13, type = c("not_applicable", "new", "same", "same", "not_applicable", "new", "new", "new", "same", "same", "not_applicable", "new", "same"))
需求说明
需要实现以下索引分配规则:
- 所有
not_applicable类型的行,索引设为0 - 每个
new类型的行,获取递增的连续索引(从1开始,依次为1、2、3...) - 每个
same类型的行,继承前一个new行的索引 - 序列始终以
new开头,若长度大于1则后续跟same
期望输出
data.frame(id = 1:13, type = c("not_applicable", "new", "same", "same", "not_applicable", "new", "new", "new", "same", "same", "not_applicable", "new", "same"), index = c(0, 1, 1, 1, 0, 2, 3, 4, 4, 4, 0, 5, 5))
尝试的代码及问题
尝试了以下代码:
df0 %>% mutate(index = with(rle(type != "not_applicable"), rep(cumsum(values) * values, lengths)))
得到的输出不符合预期,连续的new行被分配了同一个索引,而非递增的索引:
id type index 1 1 not_applicable 0 2 2 new 1 3 3 same 1 4 4 same 1 5 5 not_applicable 0 6 6 new 2 7 7 new 2 8 8 new 2 9 9 same 2 10 10 same 2 11 11 not_applicable 0 12 12 new 3 13 13 same 3
解决方案
可以结合dplyr和tidyr的fill函数实现需求,代码如下:
library(dplyr) library(tidyr) df0 %>% # 为每个new分配递增索引,其他类型先设为NA mutate(index = ifelse(type == "new", cumsum(type == "new"), NA)) %>% # 向下填充NA,让same继承前一个new的索引 fill(index, .direction = "down") %>% # 将not_applicable对应的index替换为0 mutate(index = replace(index, type == "not_applicable", 0))
执行后将得到符合预期的输出,其中:
cumsum(type == "new")会为每个new行生成连续递增的数值fill(index, .direction = "down")将same行的NA值替换为前一个非NA的索引值- 最后用
replace把not_applicable行的索引设为0
内容的提问来源于stack exchange,提问作者Polina B
相关产品推荐
相关产品推荐

