在R中实现字符串模糊匹配,筛选两列动物数据的特定名称
R实现两列内容匹配的操作方法
首先先构造和你示例一致的测试数据:
# 构造示例数据 ColumnA <- c("Cat", "Dog", "Bird", "Fish") ColumnB <- c("*&^Cat!167676fourlegs", "90876Bird##wings", "***&&Dog")
需求1:找出ColumnA中未被ColumnB内容包含的动物
我们可以用基础R自带的grepl()函数完成匹配判断,该函数会返回布尔值标识指定的匹配内容是否存在于目标字符串中,结合sapply()遍历ColumnA的所有动物名称即可完成批量判断:
# 批量判断每个动物是否在ColumnB中出现过,fixed=TRUE表示按纯文本匹配,避免正则特殊字符干扰 match_status <- sapply(ColumnA, function(animal) any(grepl(animal, ColumnB, fixed = TRUE))) # 提取未匹配的动物 unmatched_animal <- ColumnA[!match_status]
运行后得到的unmatched_animal结果为"Fish",和示例预期一致。
需求2:筛选ColumnA中可匹配到ColumnB内容的动物
直接基于上面得到的匹配状态取正向结果即可:
matched_animal <- ColumnA[match_status]
运行后得到的matched_animal结果为"Cat" "Dog" "Bird",符合示例预期。
如果你习惯使用tidyverse生态的工具,也可以用stringr包的str_detect()函数替代grepl(),逻辑完全一致:
# 需提前安装stringr包:install.packages("stringr") library(stringr) match_status_tidy <- sapply(ColumnA, function(animal) any(str_detect(ColumnB, fixed(animal)))) unmatched_animal_tidy <- ColumnA[!match_status_tidy] matched_animal_tidy <- ColumnA[match_status_tidy]
内容的提问来源于stack exchange,提问作者Sou
相关产品推荐
相关产品推荐

