You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中按条件筛选行:基于参与者列表过滤购物数据

问题:从购物记录中筛选指定参与者的记录

数据说明

我有两个数据框:df1存储杂货店顾客的购物记录,df2存储目标参与者列表,生成代码如下:

df1 <- data.frame(Person = sample(1:5, size=10, replace = T), Object = sample(letters[1:5], size=10, replace = T))
df2 <- data.frame(Participant = c(1, 3, 5))

示例数据如下:
df1:

PersonObject
1a
2a
1c
5d
4e
1b
2a
3b
2c
5d

df2:

Participant
1
3
5

需求

我需要创建df1的子集df1.2,只保留df1$Person与df2$Participant匹配的行,预期结果如下:
df1.2:

PersonObject
1a
1c
5d
1b
3b
5d

尝试过的无效代码

我试过以下两种写法,但因为两个向量长度不匹配,都没得到正确结果:

participant <- df2$Participant 
df1.2 <- subset(df1, Person == participant)
df1.2 <- df1  %>% filter(Person == df2$Participant)

解决方法

问题出在你用了==,这个运算符是逐元素匹配,当两边向量长度不一致时会循环补齐短的那个,导致匹配逻辑错误。你需要用%in%运算符,它会检查每个元素是否存在于目标集合中,不管两边长度是否一致。

方法1:基础R的subset函数

participant <- df2$Participant 
df1.2 <- subset(df1, Person %in% participant)

方法2:dplyr的filter函数

library(dplyr)
df1.2 <- df1 %>% filter(Person %in% df2$Participant)

方法3:merge合并(可选)

如果你想用合并的方式也能实现,注意指定匹配的列名即可;若要保留原df1的顺序,加上sort=FALSE:

df1.2 <- merge(df1, df2, by.x = "Person", by.y = "Participant", sort = FALSE)

内容的提问来源于stack exchange,提问作者Matias V.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 12:35:28