如何用正则表达式查找文本中所有重复出现的单词?
全文重复单词查找解决方案
你的原有正则\b(\w+)\s+\1\b仅能匹配相邻连续重复的单词(比如"we we"),无法处理全文分散的重复单词,以下是两种可行的解决方法:
方法一:提取单词+统计次数(推荐)
通过提取文本中所有单词,统计出现次数后筛选重复项,逻辑直观且容错性强:
import re string21 = "we read all sort of books, we read sci-fi books, historical books, advanture books and etc." # 提取所有单词并统一转小写(如需区分大小写可去掉.lower()) words = re.findall(r'\b\w+\b', string21.lower()) # 统计每个单词的出现次数 word_count = {} for word in words: word_count[word] = word_count.get(word, 0) + 1 # 筛选出现次数大于1的单词 duplicate_words = [word for word, count in word_count.items() if count > 1] print(duplicate_words) # 输出: ['we', 'read', 'books']
方法二:正则正向预查(纯正则实现)
利用正向预查匹配后续存在相同单词的项,最后去重得到结果:
import re string21 = "we read all sort of books, we read sci-fi books, historical books, advanture books and etc." # 正则匹配后续存在相同单词的项 pattern = r'\b(\w+)\b(?=.*\b\1\b)' matches = re.findall(pattern, string21.lower()) # 去重得到最终重复单词列表 duplicate_words = list(set(matches)) print(duplicate_words) # 输出: ['we', 'read', 'books'](顺序可能随机)
内容的提问来源于stack exchange,提问作者Avid Programmer
相关产品推荐
相关产品推荐

