Python中Counter取Top 5词频时遗漏同频词,该如何修正?
问题分析与修复方案
你的问题出在Counter.most_common(top_words)的行为上——它返回的是前N个单个单词,而不是前N个不同计数对应的所有单词。当多个单词共享同一个计数时(比如one和night都是12次),most_common(5)只会取前5个单词,这就导致同计数的其他单词被漏掉了。
修正后的代码
from collections import defaultdict from collections import Counter # 原始词频数据 lines = Counter({'you': 29, 'i': 24, 'my': 17, 'more': 13, 'one': 12, 'night': 12, 'go': 11, 'yeah': 10}) top_words = 5 # 按计数分组所有单词,确保同计数单词不遗漏 reversed_word = defaultdict(list) for word, count in lines.items(): reversed_word[count].append(word) # 获取从高到低排序的计数,取前top_words个计数等级 sorted_counts = sorted(reversed_word.keys(), reverse=True)[:top_words] # 遍历目标计数等级,输出整理后的结果 for key in sorted_counts: print("The following words appeared {} times each: {}".format(key, ', '.join(sorted(reversed_word[key]))))
修复逻辑说明
- 先遍历所有单词完成计数分组:直接基于原始词频数据分组,不会因为取前N个单词而漏掉同计数的其他词汇;
- 锁定目标计数等级:对所有计数从高到低排序后,取前
top_words个计数区间(比如top_words=5时,对应29、24、17、13、12这五个计数); - 输出完整结果:每个计数等级下的所有单词都会被展示,自然包含了同计数的
one和night。
测试输出(top_words=5时)
The following words appeared 29 times each: you The following words appeared 24 times each: i The following words appeared 17 times each: my The following words appeared 13 times each: more The following words appeared 12 times each: night, one
内容的提问来源于stack exchange,提问作者ggchc
相关产品推荐
相关产品推荐

