Python构建嵌套字典记录单词所在文档及对应出现位置的方法
Python 嵌套字典统计单词所属文档及位置实现方案
你原有代码存在三个核心问题:
words是存储每个文件分词结果的二维列表,直接遍历item in words拿到的是单文件分词子列表,不是单个目标单词- 没有获取并存储单词在对应文件内的下标位置,且逻辑上没有做同一文件内的位置聚合
- 额外提醒:不要使用
dict作为变量名,它是Python内置类型关键字,覆盖后会影响字典原生功能使用
直接符合输出要求的实现代码
dictionary = {} textfile_list = ['file1.txt', 'file2.txt', 'file3.txt'] file_contents = ['mario luigi friend mushroom', 'rick mario morty portal summer mario', 'peter griffin shop'] words = [['mario', 'luigi', 'friend', 'mushroom'], ['rick', 'mario', 'morty', 'portal', 'summer', 'mario'], ['peter', 'griffin', 'shop']] for file_idx, filename in enumerate(textfile_list): # 取当前文件对应的分词列表 current_file_words = words[file_idx] # 遍历获取每个单词的下标和内容 for word_pos, word in enumerate(current_file_words): if word not in dictionary: dictionary[word] = [] # 检查该单词是否已有当前文件的记录 has_file_record = False for record in dictionary[word]: if filename in record: record[filename].append(word_pos) has_file_record = True break # 无当前文件记录则新增 if not has_file_record: dictionary[word].append({filename: [word_pos]})
执行print(dictionary['mario'])即可得到你期望的输出:[{'file1.txt': [0]}, {'file2.txt': [1, 5]}]
更高效的优化写法
如果处理的文件量较大,可以先用中间字典做存储避免重复遍历检查,最后再转换为目标格式,性能更好:
temp_dict = {} for file_idx, filename in enumerate(textfile_list): current_file_words = words[file_idx] for word_pos, word in enumerate(current_file_words): if word not in temp_dict: temp_dict[word] = {} if filename not in temp_dict[word]: temp_dict[word][filename] = [] temp_dict[word][filename].append(word_pos) # 转换为要求的输出结构 dictionary = { word: [{file: pos_list} for file, pos_list in file_info.items()] for word, file_info in temp_dict.items() }
内容的提问来源于stack exchange,提问作者David R
相关产品推荐
相关产品推荐

