You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中匹配含重复单词(相邻/非相邻)的正则优化问题

Solution for Precise Word Repeat Matching in R DataFrames

Let's fix your regex issues to get accurate matches for both word repeats (1+ times) and repeats of 3+ times. Your original regex only handles adjacent repeated words with trailing spaces, which is why you're seeing false positives and misses.

1. Match Sentences with Any Word Repeated 1+ Times

To catch any word (including contractions like it's) that appears at least twice, regardless of position, use this regex with Perl-compatible syntax (more reliable for backreferences):

# Filter rows where any word repeats at least once
repeat_1plus <- df[grepl("\\b([\\w']+)\\b.*\\b\\1\\b", df$Turn, perl = TRUE),]

What this regex does:

  • \\b: Ensures we match whole words (avoids partial matches like your vs yourself).
  • ([\\w']+): Captures a word including apostrophes (handles contractions like don't or it's).
  • .*: Matches any characters between the two instances of the word.
  • \\b\\1\\b: Re-matches the exact captured word as a whole word.

Fixes your original issues:

  • No more false positives: yourself and your are treated as distinct words, so the sentence "yourself , everybody 'd be changing your hair in n it ?" won't be matched.
  • Catches non-adjacent repeats: The sentence "no it 's not mother theresa , it 's saint theresa ." will be matched because it and theresa each appear twice.

2. Match Sentences with Any Word Repeated 3+ Times

To find words that appear at least three times (adjacent or non-adjacent), adjust the regex to check for two additional occurrences of the captured word:

# Filter rows where any word repeats at least three times
repeat_3plus <- df[grepl("\\b([\\w']+)\\b(.*\\b\\1\\b){2,}", df$Turn, perl = TRUE),]

What this regex does:

  • \\b([\\w']+)\\b: Captures the target word (same as before).
  • (.*\\b\\1\\b){2,}: Requires the captured word to appear at least two more times (so total 3+ occurrences), with any characters in between.

Fixes your missing match:

  • Now captures "it had steve martin in it , it 's a comedy" because it appears three times, along with "no no no i 'm dave and you 're alan ." where no is repeated three times adjacently.

Example Results

For your sample dataframe:

  • repeat_1plus will include all rows except:
    • "well this is what the grumble about do n't they ?"
    • "well where have the middle sized soda stream bottle gone ?"
    • "cos he finishes at four o'clock on that day anyway ."
  • repeat_3plus will include:
    • "no no no i 'm dave and you 're alan ."
    • "it had steve martin in it , it 's a comedy"

内容的提问来源于stack exchange,提问作者Chris Ruehlemann

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 20:07:28