You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python函数查找句子中重复单词的所有出现位置及问题排查

解决Python查找重复单词所有位置的问题

你遇到的问题大概率是因为只用了str.find()这类只返回首次匹配位置的方法,或者循环查找时没有更新起始索引,导致每次都从字符串开头找,只能拿到第一个位置。下面是完整的解决方案:

实现步骤与代码

import re
from collections import defaultdict

def find_duplicate_word_positions():
    # 获取用户输入的句子
    sentence = input("请输入要分析的句子:")
    
    # 1. 预处理:分割单词,同时处理标点和大小写
    # 用正则匹配单词(包含字母、撇号,比如don't),并转小写
    words = re.findall(r"\b\w+(?:['’]\w+)?\b", sentence.lower())
    word_count = defaultdict(int)
    
    # 统计每个单词的出现次数
    for word in words:
        word_count[word] += 1
    
    # 筛选出出现次数大于1的单词
    target_words = [word for word, count in word_count.items() if count > 1]
    
    if not target_words:
        print("没有出现次数大于1的单词")
        return
    
    # 2. 查找每个目标单词的所有位置
    result = defaultdict(list)
    lower_sentence = sentence.lower()
    
    for word in target_words:
        word_len = len(word)
        start_idx = 0
        # 循环查找所有匹配位置
        while True:
            # 从start_idx开始查找
            pos = lower_sentence.find(word, start_idx)
            if pos == -1:
                break
            # 记录位置(这里的位置是原字符串中的起始索引)
            result[word].append(pos)
            # 更新起始索引,避免重复匹配同一个位置
            start_idx = pos + word_len
    
    # 3. 输出结果
    print("出现次数大于1的单词及其位置:")
    for word, positions in result.items():
        print(f"单词 '{word}' 的位置:{positions}")

# 调用函数
find_duplicate_word_positions()

关键细节说明

  • 标点与大小写处理:用正则r"\b\w+(?:['’]\w+)?\b"可以匹配带撇号的单词(比如don't),同时转成小写避免Hello和hello被当成不同单词。
  • 循环查找位置:每次找到一个位置后,把下一次的起始索引设为pos + word_len,这样就能跳过当前匹配的单词,继续找下一个,不会重复获取同一个位置。
  • 统计次数:用collections.defaultdict统计单词出现次数,筛选出需要处理的目标单词,避免无意义的查找。

测试示例

输入句子:Hello hello, world! Hello again world.
输出结果:

出现次数大于1的单词及其位置:
单词 'hello' 的位置:[0, 7, 21]
单词 'world' 的位置:[14, 32]

内容的提问来源于stack exchange,提问作者Gian carlo Khalil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 06:50:40