R语言中按列搜索多关键词并生成对应布尔列的方法
为每个关键词生成独立布尔匹配列的解决方案
原始数据
stringstosearch <- c("to", "and", "at", "from", "is", "of") set.seed(199) datatxt <- data.frame(id = c(rnorm(5)), x = c("Contrary to popular belief, Lorem Ipsum is not simply random text.", "A Latin professor at Hampden-Sydney College in Virginia", "It has roots in a piece of classical Latin ", "literature from 45 BC, making it over 2000 years old.", "The standard chunk of Lorem Ipsum used since"))
需求说明
需要在datatxt的x列中搜索stringstosearch中的每个关键词,为每个关键词生成独立的布尔值列(包含关键词则为TRUE,否则为FALSE)。
解决方案
方法1:使用dplyr + stringr(tidyverse风格)
借助dplyr的across函数,可直接在原数据框中批量生成匹配列:
library(dplyr) library(stringr) datatxt_result <- datatxt %>% mutate(across(all_of(stringstosearch), ~str_detect(x, fixed(.x))))
all_of(stringstosearch)指定新列名直接使用关键词本身fixed(.x)确保精确匹配关键词,避免正则表达式的意外匹配(比如"to"不会匹配"tomorrow"中的子串)str_detect(x, fixed(.x))检测每行x列是否包含当前关键词
方法2:使用purrr + stringr
若需单独生成匹配列再合并,可用purrr的map_dfc批量处理:
library(purrr) library(stringr) library(tibble) # 生成所有匹配列 match_cols <- map_dfc(stringstosearch, ~tibble(!!.x := str_detect(datatxt$x, fixed(.x)))) # 合并到原数据框 datatxt_result <- bind_cols(datatxt, match_cols)
方法3:基础R实现(无需额外包)
不想加载tidyverse包的话,用基础R的sapply也能完成:
# 生成匹配矩阵,每行对应原数据的一行,每列对应一个关键词 match_matrix <- sapply(stringstosearch, function(key) grepl(fixed(key), datatxt$x)) # 合并到原数据框 datatxt_result <- cbind(datatxt, as.data.frame(match_matrix))
结果说明
生成的datatxt_result包含原数据的id、x列,以及to、and、at等关键词对应的布尔列。例如第一行的to和is列会返回TRUE,其余关键词列返回FALSE。
内容的提问来源于stack exchange,提问作者JontroPothon
相关产品推荐
相关产品推荐

