You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DataFrame列的列表中匹配值并生成标记列

解决Tidyverse中逐行匹配单个值与列表值的问题

需求说明

在R语言tidyverse环境下,需为数据框新增一列match_found:检查每行的filedate单个值是否存在于对应行的filedate_list列表中,存在则标记1,否则标记0。

原始数据框

library(tidyverse)

df_original <- tribble(
  ~record_num, ~filedate, ~filedate_list, 
  1, 1998, c(1998, 1999, 2000, 2001),
  2, 1999, c(1998, 1999, 2000, 2001),
  3, 2005, c(1998, 1999, 2000, 2001),
  4, 2006, c(1998, 1999, 2000, 2001),
)

期望结果

df_solution <- tribble(
  ~record_num, ~filedate, ~filedate_list, ~match_found,
  1, 1998, c(1998, 1999, 2000, 2001), 1, 
  2, 1999, c(1998, 1999, 2000, 2001), 1, 
  3, 2005, c(1998, 1999, 2000, 2001), 0,
  4, 2006, c(1998, 1999, 2000, 2001), 0
)

错误解法及原因

之前尝试的代码返回全0,问题出在%in%的作用逻辑:它是对整个向量做匹配,而非逐行处理每行的filedate和对应行的filedate_list。

incorrect_solution <- df_original %>%
  mutate(match_found = if_else(filedate %in% filedate_list, 1, 0))

正确解法

方法1:使用purrr::map2_int逐行映射

通过map2_int同时遍历filedate和filedate_list两列,对每行的两个值做匹配判断,返回整数型结果:

df_solution1 <- df_original %>%
  mutate(match_found = map2_int(filedate, filedate_list, ~if_else(.x %in% .y, 1L, 0L)))

方法2:使用rowwise()逐行分组处理

先通过rowwise()将数据框按行分组,此时%in%会针对每行的filedate和filedate_list做判断,最后用ungroup()取消分组恢复原结构:

df_solution2 <- df_original %>%
  rowwise() %>%
  mutate(match_found = as.integer(filedate %in% filedate_list)) %>%
  ungroup()

内容的提问来源于stack exchange,提问作者Mishalb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 11:10:24