You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言如何从data frame的text列提取含指定关键词的句子到新列

R 实现文本行中匹配关键词的句子提取

方案1:tidyverse 实现(推荐,逻辑清晰易调整)

依赖dplyr和stringr两个常用包,不需要复杂语法:

# 加载依赖包
library(dplyr)
library(stringr)

# 定义待匹配关键词
needles = c("first", "hope", "analyze", "happy")
# 构造匹配正则,\\b代表单词边界,避免部分匹配(比如happy不会匹配happily),不需要可删除\\b
match_pattern <- paste0("\\b", needles, "\\b", collapse = "|")

# 批量处理生成findings列
mydata <- mydata %>%
  mutate(
    findings = sapply(text, function(cur_text) {
      # 按句号+可选空格拆分单句,过滤空值
      sentences <- str_split(cur_text, "\\.\\s?", simplify = T) %>% .[. != ""]
      # 筛选包含任意关键词的句子
      matched_sentences <- sentences[str_detect(sentences, match_pattern)]
      # 拼接结果,无匹配返回NA
      ifelse(length(matched_sentences) > 0, paste0(matched_sentences, collapse = ". "), NA_character_)
    })
  )

方案2:Base R 实现(无需额外安装包)

完全用R自带函数实现,适配无网络/无权限装包的场景:

needles = c("first", "hope", "analyze", "happy")
match_pattern <- paste0("\\b", needles, "\\b", collapse = "|")

mydata$findings <- sapply(mydata$text, function(cur_text) {
  sentences <- unlist(strsplit(cur_text, "\\.\\s?"))
  sentences <- sentences[sentences != ""]
  matched_sentences <- sentences[grepl(match_pattern, sentences)]
  if(length(matched_sentences) > 0) paste0(matched_sentences, collapse = ". ") else NA
})

适配调整说明

  • 如果句子结尾包含问号、感叹号,可将拆分正则从\\.\\s?改为[.!?]\\s?,适配更多句式
  • 若需要模糊匹配(比如关键词first可匹配firstly),删除构造match_pattern时的\\b即可
  • 单条文本包含上千句的大体积数据场景,可将sapply替换为vapply指定返回值类型,提升运行效率

内容的提问来源于stack exchange,提问作者MDStat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 17:15:08