如何在DataFrame列的列表中匹配值并生成标记列
解决Tidyverse中逐行匹配单个值与列表值的问题
需求说明
在R语言tidyverse环境下,需为数据框新增一列match_found:检查每行的filedate单个值是否存在于对应行的filedate_list列表中,存在则标记1,否则标记0。
原始数据框
library(tidyverse) df_original <- tribble( ~record_num, ~filedate, ~filedate_list, 1, 1998, c(1998, 1999, 2000, 2001), 2, 1999, c(1998, 1999, 2000, 2001), 3, 2005, c(1998, 1999, 2000, 2001), 4, 2006, c(1998, 1999, 2000, 2001), )
期望结果
df_solution <- tribble( ~record_num, ~filedate, ~filedate_list, ~match_found, 1, 1998, c(1998, 1999, 2000, 2001), 1, 2, 1999, c(1998, 1999, 2000, 2001), 1, 3, 2005, c(1998, 1999, 2000, 2001), 0, 4, 2006, c(1998, 1999, 2000, 2001), 0 )
错误解法及原因
之前尝试的代码返回全0,问题出在%in%的作用逻辑:它是对整个向量做匹配,而非逐行处理每行的filedate和对应行的filedate_list。
incorrect_solution <- df_original %>% mutate(match_found = if_else(filedate %in% filedate_list, 1, 0))
正确解法
方法1:使用purrr::map2_int逐行映射
通过map2_int同时遍历filedate和filedate_list两列,对每行的两个值做匹配判断,返回整数型结果:
df_solution1 <- df_original %>% mutate(match_found = map2_int(filedate, filedate_list, ~if_else(.x %in% .y, 1L, 0L)))
方法2:使用rowwise()逐行分组处理
先通过rowwise()将数据框按行分组,此时%in%会针对每行的filedate和filedate_list做判断,最后用ungroup()取消分组恢复原结构:
df_solution2 <- df_original %>% rowwise() %>% mutate(match_found = as.integer(filedate %in% filedate_list)) %>% ungroup()
内容的提问来源于stack exchange,提问作者Mishalb
相关产品推荐
相关产品推荐

