如何用R语言按规则从多段文本提取含指定词汇的句子
提取含特定关键词的句子解决方案
需求
- 从多行文本中提取符合以下规则的句子:
- 包含单词
bonus或incentive(不区分大小写); - 句子以标点(
. ! ?)、换行或控制字符分隔。
- 包含单词
测试数据
text <- c("This is a sentence. $5k SIGN-ON BONUS offered. This is another sentence. Salary is $15.00 per hours. Another", "This is a sentence. Retention bonus of $5,000 offered! This is another sentence. Salary is $15.00 per hours? Another", "This is a sentence. $5k incentive offered! This is another sentence. Salary is $15.00 per hours. Another", "This is a sentence\n \n$5000 sign-on Bonus offered\n \nThis is another sentence\n \nSalary is $15.00 per hours\n \nAnother", "This is a sentence\n\nRetention bonus of $5000 offered\n\nThis is another sentence\n\nSalary is $15.00 per hours\n\nAnother", "This is a sentence\n \n$5k incentive offered\n \nThis is another sentence\n Salary is $15.00 per hours\nAnother", "This is a sentence. $5k signing bonus offered! This is another sentence. Salary is $15.00 per hours? Another", "This is a sentence. This is another sentence. $5k incentive offered! Salary is $15.00 per hours? Another")
之前的尝试及问题
使用stringr::str_extract的正则表达式未得到理想结果:
stringr::str_extract(text, "[[:print:]]*(?i)bonus|(?i)incentive[[:print:]]*[[:cntrl:]]|[[:punct:]]") [1] "This is a sentence. $5k SIGN-ON BONUS" "This is a sentence. Retention bonus" [3] "." "$5000 sign-on Bonus" [5] "Retention bonus" "incentive offered\n" [7] "." "."
期望输出:
[1] "$5k SIGN-ON BONUS offered" "Retention bonus of $5,000 offered" [3] "$5k incentive offered" "$5000 sign-on Bonus offered" [5] "Retention bonus of $5000 offered" "$5k incentive offered" [7] "$5k signing bonus offered" "$5k incentive offered"
解决方案
之前的正则逻辑未精准界定句子边界,导致提取结果混乱。以下两种方法可解决问题:
方法一:拆分句子后筛选(逻辑清晰)
先将文本按句子分隔符拆分为独立句子,再筛选包含目标关键词的句子:
library(stringr) # 定义句子分隔规则:标点后接空白/控制字符,或直接控制字符 split_pattern <- "(?<=[.!?])\\s*|\\s*[[:cntrl:]]\\s*" # 遍历每个文本元素,拆分、筛选、清理空白 result <- sapply(text, function(x) { sentences <- str_split(x, split_pattern)[[1]] target_sentences <- str_subset(sentences, regex("bonus|incentive", ignore_case = TRUE)) str_trim(target_sentences) }, USE.NAMES = FALSE) # 整理为一维向量 result <- unlist(result) print(result)
方法二:正则直接提取(一步到位)
调整正则表达式,用断言精准匹配句子的前后边界:
library(stringr) # 匹配包含关键词的完整句子,忽略大小写 extract_pattern <- regex("(?<=^|\\s|[.!?]|\\n)\\s*([^.!?\\n]+(bonus|incentive)[^.!?\\n]+)\\s*(?=[.!?]|\\n|$)", ignore_case = TRUE) result <- str_extract(text, extract_pattern) # 清理句子前后的空白字符 result <- str_trim(result) print(result)
输出验证
两种方法均可得到符合预期的输出:
[1] "$5k SIGN-ON BONUS offered" "Retention bonus of $5,000 offered" [3] "$5k incentive offered" "$5000 sign-on Bonus offered" [5] "Retention bonus of $5000 offered" "$5k incentive offered" [7] "$5k signing bonus offered" "$5k incentive offered"
逻辑说明
- 拆分法:先拆分再筛选,逻辑直观,避免正则边界匹配的复杂问题;
- 直接提取法:通过正向/反向断言锁定句子的起始和结束位置,确保只提取包含关键词的完整句子。
内容的提问来源于stack exchange,提问作者nebulous
相关产品推荐
相关产品推荐

