You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Counter生成的列表绘制词频排名与词频数量对应关系图

解决方案

1. 提取指定词的排名

Counter.most_common()返回的是按词频降序排列的(词, 词频)元组列表,列表的索引值+1就是对应词的排名。
如果需要频繁查询不同词的排名,建议先构造词到排名的映射字典:

from collections import Counter

# 替换为你自己的Counter对象
word_counter = Counter(white_token)
# 拿到按词频排序后的完整列表
sorted_words = word_counter.most_common()
# 构造 词->排名 的映射字典,排名从1开始计数
word_to_rank = {word: idx+1 for idx, (word, count) in enumerate(sorted_words)}

# 直接查询"to"的排名
rank_of_to = word_to_rank.get("to")
print(rank_of_to) # 输出你需要的3

如果只是单次查询某个词的排名,也可以直接遍历匹配:

target_word = "to"
for idx, (word, count) in enumerate(word_counter.most_common()):
    if word == target_word:
        print(f"{target_word}的排名是{idx+1}")
        break

2. 提取rank和count数据绘制关系图

直接遍历排序后的列表,分别提取排名和对应的词频即可绘图,词频分布通常符合幂律分布,用双对数坐标展示更直观,以下是matplotlib的实现示例:

import matplotlib.pyplot as plt

# 提取所有排名、对应词频数据
ranks = [idx+1 for idx, (word, count) in enumerate(sorted_words)]
counts = [count for idx, (word, count) in enumerate(sorted_words)]

# 绘图
plt.figure(figsize=(10,6))
plt.loglog(ranks, counts, marker='.', linestyle='')
plt.xlabel('Rank (排名)')
plt.ylabel('Count (词频)')
plt.title('词频-排名分布')
plt.grid(True)
plt.show()

补充:同频词并列排名处理

如果存在多个词词频相同的情况,上述方法会按most_common()默认的插入先后顺序给同频词不同排名,如需同频词使用相同排名,可按以下逻辑处理:

word_to_rank = {}
prev_count = None
current_rank = 1
for idx, (word, count) in enumerate(sorted_words):
    if count != prev_count:
        current_rank = idx + 1
        prev_count = count
    word_to_rank[word] = current_rank

内容的提问来源于stack exchange,提问作者bbluo22

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 21:39:00