如何用R语言grepl函数提取匹配的具体品牌元素?
问题:从匹配语句中提取对应品牌元素
已有数据框df1的品牌列:
brands 1 Nike 2 Adidas 3 D&G
数据框df2的语句列:
statements 1 I love Nike 2 I don't like Adidas 3 I hate Puma
之前使用以下代码筛选出了包含目标品牌的df2子集:
subset_df2 <- df2[grepl(paste(df1$brands, collapse="|"), ignore.case=TRUE, df2$statements), ]
现在需要提取每个匹配语句中对应的具体品牌元素(得到向量[Nike, Adidas]),而非完整语句,实现方法如下:
方法1:使用stringr包(简洁推荐)
stringr包的str_extract函数可直接从字符串中提取匹配正则表达式的内容,步骤如下:
# 构造示例数据(已有数据可跳过) df1 <- data.frame(brands = c("Nike", "Adidas", "D&G"), stringsAsFactors = FALSE) df2 <- data.frame(statements = c("I love Nike", "I don't like Adidas", "I hate Puma"), stringsAsFactors = FALSE) # 首次使用需安装包:install.packages("stringr") library(stringr) # 生成匹配所有品牌的正则模式 brand_pattern <- paste(df1$brands, collapse = "|") # 提取每个语句中匹配的品牌,忽略大小写 matched_brands <- str_extract(df2$statements, regex(brand_pattern, ignore_case = TRUE)) # 移除未匹配到品牌的NA值 matched_brands <- na.omit(matched_brands)
执行后matched_brands的结果为:
[1] "Nike" "Adidas"
方法2:基础R实现(无需额外包)
如果不想加载第三方包,可用基础R的regmatches和regexec函数实现:
# 构造示例数据(已有数据可跳过) df1 <- data.frame(brands = c("Nike", "Adidas", "D&G"), stringsAsFactors = FALSE) df2 <- data.frame(statements = c("I love Nike", "I don't like Adidas", "I hate Puma"), stringsAsFactors = FALSE) # 生成匹配模式 brand_pattern <- paste(df1$brands, collapse = "|") # 匹配并提取品牌 matches <- regmatches(df2$statements, regexec(brand_pattern, df2$statements, ignore.case = TRUE)) matched_brands <- unlist(lapply(matches, function(x) if(length(x) > 0) x[1] else NA)) # 移除NA值 matched_brands <- na.omit(matched_brands)
得到的结果与方法1一致。
内容的提问来源于stack exchange,提问作者Max Schmidt
相关产品推荐
相关产品推荐

