You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R提取正则匹配结果并合并去重后的匹配项?

解决方法

要实现提取films前的单词并去重后合并,只需在提取所有匹配项后添加去重步骤即可,具体代码如下:

基础写法

text1 <- "Netflix announced 34 new Korean films to hit the streaming platform in 2023, along with 12 Japanese films. The upcoming titles, which Netflix calls their “biggest-ever lineup of Korean films and series."

pattern <- "\\b[[:alpha:]]+\\b(?=\\sfilms)"

# 提取所有匹配的单词
matches <- str_extract_all(text1, pattern)[[1]]
# 去重
unique_matches <- unique(matches)
# 合并为目标格式
paste(unique_matches, collapse = " | ")

管道式写法(tidyverse风格)

如果你习惯用管道操作,也可以写成更简洁的形式:

library(stringr)
library(purrr)

text1 <- "Netflix announced 34 new Korean films to hit the streaming platform in 2023, along with 12 Japanese films. The upcoming titles, which Netflix calls their “biggest-ever lineup of Korean films and series."

pattern <- "\\b[[:alpha:]]+\\b(?=\\sfilms)"

str_extract_all(text1, pattern) %>%
  pluck(1) %>%  # 从列表中取出匹配的字符向量
  unique() %>%  # 去除重复项
  paste(collapse = " | ")  # 合并为指定格式的字符串

运行以上代码后,输出结果即为:

'Korean | Japanese'

关键步骤说明

  • str_extract_all(text1, pattern)[[1]]:将提取到的列表型结果转换为字符向量,方便后续处理
  • unique():对字符向量中的元素去重,保留唯一值
  • paste(..., collapse = " | "):将去重后的元素用|连接成单个字符串

内容的提问来源于stack exchange,提问作者Raj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 14:35:27