R正则正向环视:筛选特定位置含woman的DataFrame行
R DataFrame 筛选符合特定位置条件的行
直接用以下代码即可实现你的需求,排除she be woman并保留其他符合条件的行:
library(dplyr) library(stringr) # 原始数据框 phrases_with_woman <- structure(list(phrase = c("woman get degree", "woman obtain justice", "session woman vote for member", "woman have to end", "woman have no existence", "woman lose right", "woman be much", "woman mix at dance", "woman vote as member", "woman have power", "woman act only", "she be woman", "no committee woman passed vote")), row.names = c(NA, -13L), class = "data.frame") # 筛选逻辑 matches <- phrases_with_woman %>% filter(str_detect(phrase, "^woman|^\\w+\\swoman|^(no|not|never)\\s\\w+\\swoman"))
正则表达式对应条件拆解
^woman:匹配以woman开头的短语,对应条件1(woman是句中第一个词)。^\\w+\\swoman:匹配首词为任意单词、第二个词是woman的短语,对应条件2(woman是句中第二个词)。^(no|not|never)\\s\\w+\\swoman:匹配首词为no/not/never、第三个词是woman的短语,对应条件3(woman是句中第三个词,且前面的指定否定词位于句首)。
原代码问题分析
你之前使用的(?<=woman\\s)\\w+是正向后顾断言,作用是匹配woman之后的单词,这和你需要的“按位置筛选woman”的逻辑完全不符,因此会错误匹配大量不符合要求的行。
内容的提问来源于stack exchange,提问作者generic
相关产品推荐
相关产品推荐

