You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将字符串中每个单词的起止索引映射为字典的问题

问题解决:正确生成单词索引并转换为字典

问题出在哪

你现在得到重复索引的根本原因,不是defaultdict的问题,而是你提供的boundaries_list本身就有错误——里面重复出现的单词(比如'a'、'of'),对应的索引全都是它们首次出现的位置,根本不是后续实际出现的索引值。

第一步:正确生成单词的起止索引(符合要求:索引从1开始,省略空格计数)

要准确获取每个单词的起止索引(跳过空格,只给有效字符编连续索引),可以用下面的代码实现:

text = "i have a list of lists that contain a word and there indices my method works except with repeated words like of or a or the or it"

word_boundaries = []
current_index = 1  # 索引从1开始

for char in text:
    if char == ' ':
        continue  # 跳过空格,不计数
    # 检查当前字符是不是单词的开头
    if (current_index == 1) or (text[current_index - 2] == ' '):
        # 找到单词起始位置,现在找单词结束位置
        word_end = current_index
        # 向后遍历直到遇到空格或字符串结束
        while word_end <= len(text) and text[word_end - 1] != ' ':
            word_end += 1
        word_end -= 1  # 回退到单词最后一个字符的索引
        word = text[current_index - 1 : word_end - 1 + 1]  # 从原始字符串截取单词
        word_boundaries.append([word, [current_index, word_end]])
        # 直接跳到单词结束后的位置,避免重复遍历
        current_index = word_end + 1
    else:
        current_index += 1

这段代码会严格按照“省略空格、索引从1开始”的规则,生成每个单词对应的真实起止索引,比如三次出现的'a'会分别对应[3,3]、[27,27]、[76,76](可自行运行验证)。

第二步:转换为字典(保留所有重复单词的索引)

有了正确的word_boundaries列表后,转换为字典的代码可以直接简化,不用绕弯子生成小字典再合并:

方法1:用defaultdict

from collections import defaultdict

result_dict = defaultdict(list)
for word, indices in word_boundaries:
    result_dict[word].append(indices)

方法2:用普通字典(无需导入模块)

result_dict = {}
for word, indices in word_boundaries:
    if word not in result_dict:
        result_dict[word] = []
    result_dict[word].append(indices)

两种方法都能得到正确结果:每个单词对应的列表里,会按出现顺序保存所有真实的起止索引,不会出现重复的首次索引。

示例结果

比如'a'在字典中的值会是:

'a': [[3, 3], [27, 27], [76, 76]]

完全符合每个单词实际出现的位置。


内容的提问来源于stack exchange,提问作者CodeDependency

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 17:20:35