如何用Python函数查找句子中重复单词的所有出现位置及问题排查
解决Python查找重复单词所有位置的问题
你遇到的问题大概率是因为只用了str.find()这类只返回首次匹配位置的方法,或者循环查找时没有更新起始索引,导致每次都从字符串开头找,只能拿到第一个位置。下面是完整的解决方案:
实现步骤与代码
import re from collections import defaultdict def find_duplicate_word_positions(): # 获取用户输入的句子 sentence = input("请输入要分析的句子:") # 1. 预处理:分割单词,同时处理标点和大小写 # 用正则匹配单词(包含字母、撇号,比如don't),并转小写 words = re.findall(r"\b\w+(?:['’]\w+)?\b", sentence.lower()) word_count = defaultdict(int) # 统计每个单词的出现次数 for word in words: word_count[word] += 1 # 筛选出出现次数大于1的单词 target_words = [word for word, count in word_count.items() if count > 1] if not target_words: print("没有出现次数大于1的单词") return # 2. 查找每个目标单词的所有位置 result = defaultdict(list) lower_sentence = sentence.lower() for word in target_words: word_len = len(word) start_idx = 0 # 循环查找所有匹配位置 while True: # 从start_idx开始查找 pos = lower_sentence.find(word, start_idx) if pos == -1: break # 记录位置(这里的位置是原字符串中的起始索引) result[word].append(pos) # 更新起始索引,避免重复匹配同一个位置 start_idx = pos + word_len # 3. 输出结果 print("出现次数大于1的单词及其位置:") for word, positions in result.items(): print(f"单词 '{word}' 的位置:{positions}") # 调用函数 find_duplicate_word_positions()
关键细节说明
- 标点与大小写处理:用正则
r"\b\w+(?:['’]\w+)?\b"可以匹配带撇号的单词(比如don't),同时转成小写避免Hello和hello被当成不同单词。 - 循环查找位置:每次找到一个位置后,把下一次的起始索引设为
pos + word_len,这样就能跳过当前匹配的单词,继续找下一个,不会重复获取同一个位置。 - 统计次数:用
collections.defaultdict统计单词出现次数,筛选出需要处理的目标单词,避免无意义的查找。
测试示例
输入句子:Hello hello, world! Hello again world.
输出结果:
出现次数大于1的单词及其位置: 单词 'hello' 的位置:[0, 7, 21] 单词 'world' 的位置:[14, 32]
内容的提问来源于stack exchange,提问作者Gian carlo Khalil
相关产品推荐
相关产品推荐

