R语言标记两列交叉值在各列首次出现的实现方法
实现思路
先按行的先后顺序,记录所有值第一次出现的对应行号,再分别匹配x、y列的每个值,判断当前行号是否等于该值首次出现的行号即可。
完整代码
library(dplyr) # 构造示例数据 df <- tibble(x = c(1, 2, 2, 3, 7), y = c(7, 7, 8, 9, 10)) # 生成按行交错的所有值序列,统计每个值首次出现的行号 all_vals <- c(rbind(df$x, df$y)) first_occur_row <- tibble(val = all_vals) %>% mutate(row = (row_number() + 1) %/% 2) %>% group_by(val) %>% slice_min(row, n = 1) %>% ungroup() %>% deframe() # 新增标记列 df <- df %>% mutate( first_occurance_x = row_number() == first_occur_row[as.character(x)], first_occurance_y = row_number() == first_occur_row[as.character(y)] )
输出结果
print(df) #> # A tibble: 5 × 4 #> x y first_occurance_x first_occurance_y #> <dbl> <dbl> <lgl> <lgl> #> 1 1 7 TRUE TRUE #> 2 2 7 TRUE FALSE #> 3 2 8 FALSE TRUE #> 4 3 9 TRUE TRUE #> 5 7 10 FALSE TRUE
如果你需要输出的是0/1的数值型标记而非逻辑值,只需要在mutate时把判断结果用as.integer()包裹即可。
内容的提问来源于stack exchange,提问作者John-Henry
相关产品推荐
相关产品推荐

