Python实现Reverse Index疑问:拆分单词的思路是否正确?
反向索引函数实现思路与正确代码
你的思路完全正确——先拆分每个字符串为单词,再通过循环构建关键词到索引列表的映射,问题大概率出在细节处理上。
常见的出错细节:
- 未将所有单词统一转为小写,导致同一单词的大小写形式被识别为不同键(比如"Hello"和"hello")
- 循环时未正确记录字符串的索引,或者同一个字符串内重复出现的单词重复添加了相同索引
- 拆分单词时未处理空格以外的分隔符(不过题目未明确要求的话,默认按空格拆分即可)
正确实现代码:
def reverse_index(strings): index_dict = {} # 遍历每个字符串,同时记录其索引 for idx, content in enumerate(strings): # 转小写后拆分单词 words = content.lower().split() # 去重当前字符串内的单词,避免重复添加同一索引 unique_words = set(words) for word in unique_words: # 单词不在字典中则初始化空列表 if word not in index_dict: index_dict[word] = [] # 添加当前字符串的索引 index_dict[word].append(idx) # 可选:对索引列表排序,保证输出顺序统一 for key in index_dict: index_dict[key].sort() return index_dict
测试示例:
输入:
sample_input = [ "Hello world", "World of programming", "Hello programming" ] print(reverse_index(sample_input))
输出:
{ 'hello': [0, 2], 'world': [0, 1], 'of': [1], 'programming': [1, 2] }
内容的提问来源于stack exchange,提问作者Jerry Cohen
相关产品推荐
相关产品推荐

