如何用dplyr结合正则匹配符合特定规则的info文本行?
解决方案
你需要调整正则表达式来覆盖所有三个条件,以下是修正后的代码:
library(dplyr) library(stringr) df_new <- data.frame( text=c('info is given','he is given info. in the class', 'she needs info2','why not having information', 'his info# missing', 'info12 and packages are given', 'parainfo is ready','info. was awarded', 'meeting is with .info')) df_result <- df_new %>% mutate(text = tolower(text)) %>% mutate(strings_detected = as.integer(str_detect(text, "(^|\\s)info(?:\\s|$|\\W|\\d+)"))) # 查看结果 print(df_result)
正则表达式说明
(^|\\s)info(?:\\s|$|\\W|\\d+) 各部分含义:
(^|\\s):匹配字符串开头或空白字符,确保info前面不是其他字符(排除.info、parainfo这类情况)info:目标匹配术语(?:\\s|$|\\W|\\d+):非捕获组,匹配以下任意一种情况:\\s:空白字符(对应条件1:info后无其他字符)$:字符串结尾(对应条件1:info后无其他字符)\\W:任意非字母数字字符(对应条件2:info后紧跟英文句号或特殊字符)\\d+:一个或多个数字(对应条件3:info后紧跟数字)
输出结果
运行代码后得到的结果与你的期望完全一致:
text strings_detected 1 info is given 1 2 he is given info. in the class 1 3 she needs info2 1 4 why not having information 0 5 his info# missing 1 6 info12 and packages are given 1 7 parainfo is ready 0 8 info. was awarded 1 9 meeting is with .info 0
内容的提问来源于stack exchange,提问作者Alex
相关产品推荐
相关产品推荐

