You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则表达式查找文本中所有重复出现的单词?

全文重复单词查找解决方案

你的原有正则\b(\w+)\s+\1\b仅能匹配相邻连续重复的单词(比如"we we"),无法处理全文分散的重复单词,以下是两种可行的解决方法:

方法一:提取单词+统计次数(推荐)

通过提取文本中所有单词,统计出现次数后筛选重复项,逻辑直观且容错性强:

import re

string21 = "we read all sort of books, we read sci-fi books, historical books, advanture books and etc."
# 提取所有单词并统一转小写(如需区分大小写可去掉.lower())
words = re.findall(r'\b\w+\b', string21.lower())
# 统计每个单词的出现次数
word_count = {}
for word in words:
    word_count[word] = word_count.get(word, 0) + 1
# 筛选出现次数大于1的单词
duplicate_words = [word for word, count in word_count.items() if count > 1]
print(duplicate_words)  # 输出: ['we', 'read', 'books']

方法二:正则正向预查(纯正则实现)

利用正向预查匹配后续存在相同单词的项,最后去重得到结果:

import re

string21 = "we read all sort of books, we read sci-fi books, historical books, advanture books and etc."
# 正则匹配后续存在相同单词的项
pattern = r'\b(\w+)\b(?=.*\b\1\b)'
matches = re.findall(pattern, string21.lower())
# 去重得到最终重复单词列表
duplicate_words = list(set(matches))
print(duplicate_words)  # 输出: ['we', 'read', 'books'](顺序可能随机)

内容的提问来源于stack exchange,提问作者Avid Programmer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 19:27:55