Python嵌套字典迭代异常:仅遍历列表最后元素问题排查
问题排查:TF-IDF计算仅保留文档最后一个词的问题
问题场景
我有一个以文档ID为键、对应词列表为值的字典data_docs(共包含5个文档),希望对每个文档执行TF-IDF计算后,将词与计算结果存入嵌套字典中。但迭代后出现异常:最终生成的嵌套字典仅保留了每个词列表的最后一个元素。
原字典data_docs
{'doc01': ['simpl', 'hello', 'world', 'test', 'python', 'code'], 'doc02': ['today', 'wonder', 'day'], 'doc03': ['studi', 'pac', 'today'], 'doc04': ['write', 'need', 'cup', 'coffe'], 'doc05': ['finish', 'pac', 'use', 'python']}
当前错误结果
{'doc01': {'code': 0.6989700043360189}, 'doc02': {'day': 0.6989700043360189}, 'doc03': {'today': 0.3979400086720376}, 'doc04': {'coffe': 0.6989700043360189}, 'doc05': {'python': 0.3979400086720376}}
错误代码
def tfidf (data, idf_score): #function, 2 dictionaries as parameters tfidf = {} #dict for output for word, val in data.items(): #for each word and value in data_docs(first dict) for v in val: #for each value in each list a = val.count(v) #count the number of times that appears in that list scores = {v :a * idf_score[v]} # dictionary that will act as value in the nested tfidf[word] = scores #final dictionary, the key is doc01,doc02... and the value the above dict return tfidf tfidf(data_docs, idf_score)
错误原因
核心问题是每次内层循环都重新赋值了scores字典,而非在每个文档的循环周期内先初始化空字典再逐个添加键值对:
- 外层循环遍历文档时,没有为当前文档创建专属的空
scores字典 - 内层循环每处理一个词,就会新建一个仅包含当前词的
scores字典,直接覆盖之前的内容 - 内层循环结束后,
scores仅保留最后一个词的计算结果,最终赋值给文档ID时自然只显示最后一个词
修正后的代码
需在每个文档的循环开头初始化scores为空字典,再通过键值对赋值的方式逐个添加计算结果:
def tfidf(data, idf_score): tfidf = {} # 遍历每个文档,变量名更清晰:doc_id是文档ID,word_list是对应词列表 for doc_id, word_list in data.items(): scores = {} # 为当前文档初始化空字典,存储所有词的TF-IDF值 for word in word_list: tf = word_list.count(word) scores[word] = tf * idf_score[word] tfidf[doc_id] = scores return tfidf # 调用函数 result = tfidf(data_docs, idf_score)
优化建议(可选)
原代码中word_list.count(word)每次遍历都要扫描整个列表,长列表场景下效率较低。可以用collections.Counter先一次性统计所有词的出现次数,提升性能:
from collections import Counter def tfidf(data, idf_score): tfidf = {} for doc_id, word_list in data.items(): # 一次性统计当前文档所有词的出现次数 tf_counter = Counter(word_list) # 用字典推导式生成当前文档的TF-IDF结果 scores = {word: tf_counter[word] * idf_score[word] for word in word_list} tfidf[doc_id] = scores return tfidf
内容的提问来源于stack exchange,提问作者Josepiedehierroa
相关产品推荐
相关产品推荐

