R语言:检测tibble中句子是否包含指定姓名列表元素
解决方案:判断句子是否包含姓名列表中的任意姓名
要实现需求,核心是把姓名列表合并成一个正则匹配模式,让每个句子一次性检查是否包含任意姓名。以下是两种可行方案:
方案一:使用tidyverse工具链
library(tidyverse) # 原始数据 df <- tibble(sentences = c("Bob is looking for something", "Adriana has an umbrella", "Michael is looking at...")) names <- tibble(names = c("Bob", "Mary", "Michael", "John", "Etc.")) # 生成匹配规则:用单词边界\\b确保匹配完整姓名,避免部分匹配(比如不会把"Bobbie"误判为包含"Bob") name_pattern <- str_c("\\b", names$names, "\\b", collapse = "|") # 添加check列 df <- df %>% mutate(check = str_detect(sentences, name_pattern)) # 输出结果 df
运行后会得到目标tibble:
# A tibble: 3 × 2 sentences check <chr> <lgl> 1 Bob is looking for something TRUE 2 Adriana has an umbrella FALSE 3 Michael is looking at... TRUE
方案二:使用Base R
如果不想加载tidyverse,用Base R也能实现:
# 原始数据 df <- tibble(sentences = c("Bob is looking for something", "Adriana has an umbrella", "Michael is looking at...")) names <- tibble(names = c("Bob", "Mary", "Michael", "John", "Etc.")) # 生成匹配模式 name_pattern <- paste0("\\b", names$names, "\\b", collapse = "|") # 新增check列 df$check <- grepl(name_pattern, df$sentences, perl = TRUE)
分析你之前的错误
- 方法一:
grepl(pattern = names$names, x = df$sentences)的问题在于,当pattern是向量时,grepl会按循环规则匹配(第一个姓名匹配第一个句子,第二个姓名匹配第二个句子……),而非每个句子匹配所有姓名,所以结果不符合预期。 - 方法二:逻辑完全错误:
names$names %in% df$sentences是判断姓名是否和整个句子完全相等,而非句子包含姓名;同时str_detect的参数顺序也用反了,它的第一个参数是待检查的字符串,第二个才是匹配模式。
内容的提问来源于stack exchange,提问作者Malichot
相关产品推荐
相关产品推荐

