如何使用Counter生成的列表绘制词频排名与词频数量对应关系图
解决方案
1. 提取指定词的排名
Counter.most_common()返回的是按词频降序排列的(词, 词频)元组列表,列表的索引值+1就是对应词的排名。
如果需要频繁查询不同词的排名,建议先构造词到排名的映射字典:
from collections import Counter # 替换为你自己的Counter对象 word_counter = Counter(white_token) # 拿到按词频排序后的完整列表 sorted_words = word_counter.most_common() # 构造 词->排名 的映射字典,排名从1开始计数 word_to_rank = {word: idx+1 for idx, (word, count) in enumerate(sorted_words)} # 直接查询"to"的排名 rank_of_to = word_to_rank.get("to") print(rank_of_to) # 输出你需要的3
如果只是单次查询某个词的排名,也可以直接遍历匹配:
target_word = "to" for idx, (word, count) in enumerate(word_counter.most_common()): if word == target_word: print(f"{target_word}的排名是{idx+1}") break
2. 提取rank和count数据绘制关系图
直接遍历排序后的列表,分别提取排名和对应的词频即可绘图,词频分布通常符合幂律分布,用双对数坐标展示更直观,以下是matplotlib的实现示例:
import matplotlib.pyplot as plt # 提取所有排名、对应词频数据 ranks = [idx+1 for idx, (word, count) in enumerate(sorted_words)] counts = [count for idx, (word, count) in enumerate(sorted_words)] # 绘图 plt.figure(figsize=(10,6)) plt.loglog(ranks, counts, marker='.', linestyle='') plt.xlabel('Rank (排名)') plt.ylabel('Count (词频)') plt.title('词频-排名分布') plt.grid(True) plt.show()
补充:同频词并列排名处理
如果存在多个词词频相同的情况,上述方法会按most_common()默认的插入先后顺序给同频词不同排名,如需同频词使用相同排名,可按以下逻辑处理:
word_to_rank = {} prev_count = None current_rank = 1 for idx, (word, count) in enumerate(sorted_words): if count != prev_count: current_rank = idx + 1 prev_count = count word_to_rank[word] = current_rank
内容的提问来源于stack exchange,提问作者bbluo22
相关产品推荐
相关产品推荐

