如何用grepl多数字条件过滤R数据框中的目标行
R语言数据框多数字条件精准筛选
原始数据
创建数据框的代码:
df <- data.frame( player = c('Player To Have 1 Or More Shots On Target', 'Player To Have 1 Or More Shots On Target', 'Player To Have 2 Or More Shots On Target', 'Player To Have 3 Or More Shots On Target', 'Player To Have 1 Or More Shots On Target in 1st Half'))
数据框输出结果:
player 1 Player To Have 1 Or More Shots On Target 2 Player To Have 1 Or More Shots On Target 3 Player To Have 2 Or More Shots On Target 4 Player To Have 3 Or More Shots On Target 5 Player To Have 1 Or More Shots On Target in 1st Half
需求
仅筛选符合Player To Have X Or More Shots On Target格式的行(X为1、2、3、4等任意正整数),排除带有额外内容(如第5行的in 1st Half)的行,最终保留前4行。
现有尝试及报错
已实现仅匹配数字1的筛选代码:
df2 <- dplyr::filter(df, grepl("Player To Have 1 Or More Shots On Target", player))
尝试匹配多个数字时出现报错,代码如下:
number_of_shots <- c("1","2") df2 <- dplyr::filter(df, grepl("Player To Have", number_of_shots, "Or More Shots On Target", player))
报错信息:
Error in `dplyr::filter()`: ℹ In argument: `grepl(...)`. Caused by error: ! `..1` must be of size 5 or 1, not size 2.
解决方案
方法1:匹配任意正整数格式
用正则表达式精准匹配目标格式,通过^和$确保字符串没有额外内容:
library(dplyr) df2 <- filter(df, grepl("^Player To Have \\d+ Or More Shots On Target$", player))
^:标记字符串开头\\d+:匹配1个或多个数字(对应X)$:标记字符串结尾,彻底排除带额外内容的行
方法2:指定匹配特定数字集合
如果只需要匹配指定的几个数字(比如1、2、3),可以构造精准匹配模式:
library(dplyr) target_numbers <- c("1", "2", "3") # 拼接正则模式,匹配指定数字 pattern <- paste0("^Player To Have (", paste(target_numbers, collapse = "|"), ") Or More Shots On Target$") df2 <- filter(df, grepl(pattern, player))
这种方式可灵活控制需要匹配的数字范围,同时保证格式精准。
验证结果
运行上述代码后,df2将保留前4行数据:
> df2 player 1 Player To Have 1 Or More Shots On Target 2 Player To Have 1 Or More Shots On Target 3 Player To Have 2 Or More Shots On Target 4 Player To Have 3 Or More Shots On Target
内容的提问来源于stack exchange,提问作者Michael
相关产品推荐
相关产品推荐

