R语言如何从data frame的text列提取含指定关键词的句子到新列
R 实现文本行中匹配关键词的句子提取
方案1:tidyverse 实现(推荐,逻辑清晰易调整)
依赖dplyr和stringr两个常用包,不需要复杂语法:
# 加载依赖包 library(dplyr) library(stringr) # 定义待匹配关键词 needles = c("first", "hope", "analyze", "happy") # 构造匹配正则,\\b代表单词边界,避免部分匹配(比如happy不会匹配happily),不需要可删除\\b match_pattern <- paste0("\\b", needles, "\\b", collapse = "|") # 批量处理生成findings列 mydata <- mydata %>% mutate( findings = sapply(text, function(cur_text) { # 按句号+可选空格拆分单句,过滤空值 sentences <- str_split(cur_text, "\\.\\s?", simplify = T) %>% .[. != ""] # 筛选包含任意关键词的句子 matched_sentences <- sentences[str_detect(sentences, match_pattern)] # 拼接结果,无匹配返回NA ifelse(length(matched_sentences) > 0, paste0(matched_sentences, collapse = ". "), NA_character_) }) )
方案2:Base R 实现(无需额外安装包)
完全用R自带函数实现,适配无网络/无权限装包的场景:
needles = c("first", "hope", "analyze", "happy") match_pattern <- paste0("\\b", needles, "\\b", collapse = "|") mydata$findings <- sapply(mydata$text, function(cur_text) { sentences <- unlist(strsplit(cur_text, "\\.\\s?")) sentences <- sentences[sentences != ""] matched_sentences <- sentences[grepl(match_pattern, sentences)] if(length(matched_sentences) > 0) paste0(matched_sentences, collapse = ". ") else NA })
适配调整说明
- 如果句子结尾包含问号、感叹号,可将拆分正则从
\\.\\s?改为[.!?]\\s?,适配更多句式 - 若需要模糊匹配(比如关键词first可匹配firstly),删除构造
match_pattern时的\\b即可 - 单条文本包含上千句的大体积数据场景,可将
sapply替换为vapply指定返回值类型,提升运行效率
内容的提问来源于stack exchange,提问作者MDStat
相关产品推荐
相关产品推荐

