如何用R提取正则匹配结果并合并去重后的匹配项?
解决方法
要实现提取films前的单词并去重后合并,只需在提取所有匹配项后添加去重步骤即可,具体代码如下:
基础写法
text1 <- "Netflix announced 34 new Korean films to hit the streaming platform in 2023, along with 12 Japanese films. The upcoming titles, which Netflix calls their “biggest-ever lineup of Korean films and series." pattern <- "\\b[[:alpha:]]+\\b(?=\\sfilms)" # 提取所有匹配的单词 matches <- str_extract_all(text1, pattern)[[1]] # 去重 unique_matches <- unique(matches) # 合并为目标格式 paste(unique_matches, collapse = " | ")
管道式写法(tidyverse风格)
如果你习惯用管道操作,也可以写成更简洁的形式:
library(stringr) library(purrr) text1 <- "Netflix announced 34 new Korean films to hit the streaming platform in 2023, along with 12 Japanese films. The upcoming titles, which Netflix calls their “biggest-ever lineup of Korean films and series." pattern <- "\\b[[:alpha:]]+\\b(?=\\sfilms)" str_extract_all(text1, pattern) %>% pluck(1) %>% # 从列表中取出匹配的字符向量 unique() %>% # 去除重复项 paste(collapse = " | ") # 合并为指定格式的字符串
运行以上代码后,输出结果即为:
'Korean | Japanese'
关键步骤说明
str_extract_all(text1, pattern)[[1]]:将提取到的列表型结果转换为字符向量,方便后续处理unique():对字符向量中的元素去重,保留唯一值paste(..., collapse = " | "):将去重后的元素用|连接成单个字符串
内容的提问来源于stack exchange,提问作者Raj
相关产品推荐
相关产品推荐

