R语言中基于滞后值过滤数据的tidy/purrr实现方案
问题需求
现有数据框df,每行对应一组item1与item2的配对数据,需完成以下处理:
- 保留
df的第一行 - 后续仅保留满足前一行的
item2值等于当前行的item1值的第一行
最终输出格式参考示例output,优先采用tidyverse或purrr工具实现,也接受其他可行方案。
原始数据
df <- structure(list(item1 = c(1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 3L, 3L, 3L, 3L, 3L, 4L, 4L, 4L, 4L, 5L, 5L, 6L, 6L, 7L), item2 = c(4L, 5L, 6L, 7L, 8L, 4L, 5L, 6L, 7L, 8L, 4L, 5L, 6L, 7L, 8L, 5L, 6L, 7L, 8L, 7L, 8L, 7L, 8L, 8L)), row.names = c(NA, -24L), class = c("tbl_df", "tbl", "data.frame"))
原始数据预览:
df #> item1 item2 #> 1 1 4 #> 2 1 5 #> 3 1 6 #> 4 1 7 #> 5 1 8 #> 6 2 4 #> 7 2 5 #> 8 2 6 #> 9 2 7 #> 10 2 8 #> 11 3 4 #> 12 3 5 #> 13 3 6 #> 14 3 7 #> 15 3 8 #> 16 4 5 #> 17 4 6 #> 18 4 7 #> 19 4 8 #> 20 5 7 #> 21 5 8 #> 22 6 7 #> 23 6 8 #> 24 7 8
目标输出
output <- data.frame(item1 = c(1,4,5,7), item2 = c(4,5,7,8)) output #> item1 item2 #> 1 1 4 #> 2 4 5 #> 3 5 7 #> 4 7 8
解决方案
方法1:purrr递推实现(推荐)
利用purrr::accumulate()进行递推查找,每次以上一行的item2为匹配条件,找到下一个符合要求的第一行,直到无匹配项为止:
library(purrr) library(dplyr) # 初始化结果列表,先放入第一行 result_list <- list(df[1, ]) # 递推查找后续符合条件的行 result_list <- accumulate(seq_len(nrow(df)-1), .init = result_list, function(current, .) { last_item2 <- current[[length(current)]]$item2 # 找到第一个item1等于last_item2的行 next_row <- df %>% filter(item1 == last_item2) %>% slice(1) if(nrow(next_row) > 0) { c(current, list(next_row)) } else { current } }) # 合并为最终数据框 final_result <- bind_rows(result_list[[length(result_list)]])
运行结果:
final_result #> item1 item2 #> 1 1 4 #> 2 4 5 #> 3 5 7 #> 4 7 8
方法2:基础R循环实现
如果不依赖tidyverse工具,可用基础R循环完成相同逻辑:
final_result <- df[1, ] current_item2 <- final_result$item2 while(TRUE) { # 定位第一个匹配的行 next_row <- df[df$item1 == current_item2, ][1, ] if(nrow(next_row) == 0) break final_result <- rbind(final_result, next_row) current_item2 <- next_row$item2 }
运行结果(行号为原始数据框行号,不影响核心值):
final_result #> item1 item2 #> 1 1 4 #> 16 4 5 #> 20 5 7 #> 24 7 8
内容的提问来源于stack exchange,提问作者CyG
相关产品推荐
相关产品推荐

