You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取相对词频?基于现有Counter代码的优化需求

如何将绝对词频转换为相对词频?

嘿,要把你当前代码里的绝对词频改成相对词频其实很简单,核心思路就是先计算对应问题下的总词数,再让每个词的绝对出现次数除以这个总词数,得到的就是该词的相对频率啦。

修改后的完整代码

from collections import Counter

for question in questions:
    print('Processing ' + str(question))
    # 先统计绝对词频(和原代码一致)
    counts = Counter(dataset_final[str(question)])
    # 计算当前问题下的总词数:所有词的出现次数之和
    total_words = sum(counts.values())
    # 用字典推导式生成相对词频(可按需保留小数位数)
    relative_freq = {word: round(count / total_words, 4) for word, count in counts.items()}
    # 按相对词频从高到低排序,和原输出格式保持一致
    sorted_relative_freq = sorted(relative_freq.items(), key=lambda x: x[1], reverse=True)
    # 打印结果
    print(f"Relative frequencies: {dict(sorted_relative_freq)}")

关键步骤解释

  • 计算总词数:sum(counts.values())会把Counter里所有词的绝对出现次数加起来,得到当前问题下的总词数。
  • 生成相对词频:通过字典推导式遍历每个词和它的绝对次数,用count / total_words计算相对频率,round(..., 4)是保留4位小数,你可以根据需求调整位数(比如改成2位)。
  • 排序输出:用sorted(...)按相对频率降序排列,和你原代码的输出顺序保持一致,方便对比查看。

示例输出

Processing 1
Relative frequencies: {'would': 0.2812, 'think': 0.1875, 'patient': 0.1719, 'condition': 0.1719, 'might': 0.1562, 'increased': 0.0156}
Processing 2
Relative frequencies: {'cancer': 0.4156, 'condition': 0.2857, 'prostate': 0.2597, 'educational': 0.0129}

内容的提问来源于stack exchange,提问作者Shelina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:48:58