在R中实现非精确字符匹配并更新dataframe的section列
问题与解决方案
问题背景
现有DataFrame estimates_df,部分数据如下:
Item section 7596 5 Gal Samandoque Cacti/Accents 7597 5 Gal Purple Prickly Pear Cacti/Accents 7598 5 Gal Banana Yucca Cacti/Accents 7599 5 Gal Yucca Vine Vines 7600 5 Gal Red Three Awn Grasses 7601 3/4" Screened To Match Existing Decomposed Granite ...
另有字符向量cactus_names存储仙人掌/多肉名称:
[1]"Prickly Pear" [2]"Samandoque" [3]"Banana Yucca" ...
需要实现:不修改Item列内容,将所有Item中包含cactus_names内任意名称的行,其section列更新为"Cacti/Succulents"。此前尝试的精确匹配代码无效:
estimates_df %>% mutate(section = ifelse(cactus_names %in% Item, "Cacti/Succulents", section))
问题原因
原代码用%in%做精确匹配,要求Item完全等于cactus_names中的某一项,但实际需求是包含匹配——即Item字符串里包含cactus_names中的任意子串。
解决方案
方法1:dplyr + stringr(推荐)
用stringr::str_c将cactus_names合并为正则匹配模式,再用str_detect检查Item是否匹配:
library(dplyr) library(stringr) # 构建正则模式:匹配任意一个仙人掌名称 cactus_pattern <- str_c(cactus_names, collapse = "|") # 更新section列 estimates_df <- estimates_df %>% mutate(section = ifelse(str_detect(Item, cactus_pattern), "Cacti/Succulents", section))
方法2:基础R实现
无需额外包,用paste构建正则,结合grepl实现匹配:
# 构建正则模式 cactus_pattern <- paste(cactus_names, collapse = "|") # 更新section列 estimates_df$section <- ifelse(grepl(cactus_pattern, estimates_df$Item), "Cacti/Succulents", estimates_df$section)
特殊情况处理
如果cactus_names里包含正则特殊字符(如.、*、+等),需要先转义避免匹配异常:
# stringr包转义 cactus_pattern <- str_c(str_escape(cactus_names), collapse = "|") # 基础R转义(需加载utils包) library(utils) cactus_pattern <- paste(escapeRegex(cactus_names), collapse = "|")
内容的提问来源于stack exchange,提问作者LoveMYMAth
相关产品推荐
相关产品推荐

