在R中匹配含重复单词(相邻/非相邻)的正则优化问题
Solution for Precise Word Repeat Matching in R DataFrames
Let's fix your regex issues to get accurate matches for both word repeats (1+ times) and repeats of 3+ times. Your original regex only handles adjacent repeated words with trailing spaces, which is why you're seeing false positives and misses.
1. Match Sentences with Any Word Repeated 1+ Times
To catch any word (including contractions like it's) that appears at least twice, regardless of position, use this regex with Perl-compatible syntax (more reliable for backreferences):
# Filter rows where any word repeats at least once repeat_1plus <- df[grepl("\\b([\\w']+)\\b.*\\b\\1\\b", df$Turn, perl = TRUE),]
What this regex does:
\\b: Ensures we match whole words (avoids partial matches likeyourvsyourself).([\\w']+): Captures a word including apostrophes (handles contractions likedon'torit's)..*: Matches any characters between the two instances of the word.\\b\\1\\b: Re-matches the exact captured word as a whole word.
Fixes your original issues:
- No more false positives:
yourselfandyourare treated as distinct words, so the sentence "yourself , everybody 'd be changing your hair in n it ?" won't be matched. - Catches non-adjacent repeats: The sentence "no it 's not mother theresa , it 's saint theresa ." will be matched because
itandtheresaeach appear twice.
2. Match Sentences with Any Word Repeated 3+ Times
To find words that appear at least three times (adjacent or non-adjacent), adjust the regex to check for two additional occurrences of the captured word:
# Filter rows where any word repeats at least three times repeat_3plus <- df[grepl("\\b([\\w']+)\\b(.*\\b\\1\\b){2,}", df$Turn, perl = TRUE),]
What this regex does:
\\b([\\w']+)\\b: Captures the target word (same as before).(.*\\b\\1\\b){2,}: Requires the captured word to appear at least two more times (so total 3+ occurrences), with any characters in between.
Fixes your missing match:
- Now captures "it had steve martin in it , it 's a comedy" because
itappears three times, along with "no no no i 'm dave and you 're alan ." wherenois repeated three times adjacently.
Example Results
For your sample dataframe:
repeat_1pluswill include all rows except:- "well this is what the grumble about do n't they ?"
- "well where have the middle sized soda stream bottle gone ?"
- "cos he finishes at four o'clock on that day anyway ."
repeat_3pluswill include:- "no no no i 'm dave and you 're alan ."
- "it had steve martin in it , it 's a comedy"
内容的提问来源于stack exchange,提问作者Chris Ruehlemann
相关产品推荐
相关产品推荐

