You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于两列过滤R语言DataFrame?含ID列与待保留ID字符串列

按逗号分隔的ID字符串过滤DataFrame行

需求是根据两列过滤DataFrame的行:一列是存储单个ID的列(示例中的carb_l),另一列是用逗号分隔的待保留ID字符串(示例中的carb_list)。

示例数据准备

先构造示例数据:

library(dplyr)

mtcars2 <- mtcars %>%
  mutate(carb_l = letters[carb],  # 单个ID列
         carb_list = "c,f,h") %>% # 待保留的ID字符串(逗号分隔)
  select(-mpg, -cyl, -disp)      # 简化显示列

head(mtcars2)

运行后输出的前几行数据:

hp drat    wt  qsec vs am gear carb carb_l carb_list
1 110 3.90 2.620 16.46  0  1    4    4      d     c,f,h
2 110 3.90 2.875 17.02  0  1    4    4      d     c,f,h
3  93 3.85 2.320 18.61  1  1    4    1      a     c,f,h
4 110 3.08 3.215 19.44  1  0    3    1      a     c,f,h
5 175 3.15 3.440 17.02  0  0    3    2      b     c,f,h
6 105 2.76 3.460 20.22  1  0    3    1      a     c,f,h

过滤实现方法

核心是把逗号分隔的carb_list字符串拆分成字符向量,再用%in%判断carb_l是否在这个向量里。

如果carb_list在所有行取值相同(比如示例场景),可以先提取转成向量再过滤,效率更高:

keep_ids <- strsplit(mtcars2$carb_list[1], ",")[[1]]
mtcars2 %>% filter(carb_l %in% keep_ids)

如果carb_list每行取值不同,需要逐行处理,结合purrr工具实现:

mtcars2 %>%
  filter(purrr::map_lgl(strsplit(carb_list, ","), ~ carb_l %in% .x))

预期输出

运行上述代码后得到的结果:

> mtcars2 %>% filter(carb_l %in% c("c","f","h"))
   hp drat   wt qsec vs am gear carb carb_l carb_list
1 180 3.07 4.07 17.4  0  0    3    3      c     c,f,h
2 180 3.07 3.73 17.6  0  0    3    3      c     c,f,h
3 180 3.07 3.78 18.0  0  0    3    3      c     c,f,h
4 175 3.62 2.77 15.5  0  1    5    6      f     c,f,h
5 335 3.54 3.57 14.6  0  1    5    8      h     c,f,h

内容的提问来源于stack exchange,提问作者one

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 17:26:05