You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python无Panda/NumPy实现小规模数据集倒排索引问题求助

倒排索引代码问题排查与修正

错误原因

你的代码仅输出最后一条数据的索引,核心是缩进逻辑错误:

  • 处理分词结果的for x in wordlist:循环写在了遍历数据集的循环之外,遍历完所有数据集后wordlist仅保存了最后一句"python rules"的分词结果,所以只生成了这两个词的索引,且索引固定为最后一句的下标2
  • 额外优化点:你已经导入了collections.defaultdict,可以简化键存在性判断的逻辑,同时建议把索引字典放在函数内部,避免全局变量污染

最小改动版本(仅修正缩进,保留你原有逻辑)

from collections import defaultdict

dataset = [
    "Python time",
    "It is that TIME",
    "python rules"
 ] 

def reverse_index(dataset):
    index_dictionary = {}
    for index in range(len(dataset)):
        phrase = dataset[index]
        words = phrase.lower()
        wordlist = words.split()
        # 把分词处理循环移到数据集遍历循环内部,每处理一句就更新一次索引
        for x in wordlist:
            if x in index_dictionary.keys():
                index_dictionary[x].append(index)
            else:
                index_dictionary[x] = [index]
    return index_dictionary

print(reverse_index(dataset))

利用defaultdict优化的版本

from collections import defaultdict

dataset = [
    "Python time",
    "It is that TIME",
    "python rules"
 ] 

def reverse_index(dataset):
    # defaultdict(list)会自动给不存在的键生成空列表作为默认值
    index_dictionary = defaultdict(list)
    # enumerate可以同时拿到下标和对应元素,简化遍历写法
    for idx, phrase in enumerate(dataset):
        words = phrase.lower().split()
        for word in words:
            index_dictionary[word].append(idx)
    # 可选:如果需要输出普通字典格式,可以转成dict
    return dict(index_dictionary)

print(reverse_index(dataset))

运行输出

两个版本运行后都会得到预期结果:

{'python': [0, 2], 'time': [0, 1], 'it': [1], 'is': [1], 'that': [1], 'rules': [2]}

内容的提问来源于stack exchange,提问作者et_naej

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 11:36:03