如何获取相对词频?基于现有Counter代码的优化需求
如何将绝对词频转换为相对词频?
嘿,要把你当前代码里的绝对词频改成相对词频其实很简单,核心思路就是先计算对应问题下的总词数,再让每个词的绝对出现次数除以这个总词数,得到的就是该词的相对频率啦。
修改后的完整代码
from collections import Counter for question in questions: print('Processing ' + str(question)) # 先统计绝对词频(和原代码一致) counts = Counter(dataset_final[str(question)]) # 计算当前问题下的总词数:所有词的出现次数之和 total_words = sum(counts.values()) # 用字典推导式生成相对词频(可按需保留小数位数) relative_freq = {word: round(count / total_words, 4) for word, count in counts.items()} # 按相对词频从高到低排序,和原输出格式保持一致 sorted_relative_freq = sorted(relative_freq.items(), key=lambda x: x[1], reverse=True) # 打印结果 print(f"Relative frequencies: {dict(sorted_relative_freq)}")
关键步骤解释
- 计算总词数:
sum(counts.values())会把Counter里所有词的绝对出现次数加起来,得到当前问题下的总词数。 - 生成相对词频:通过字典推导式遍历每个词和它的绝对次数,用
count / total_words计算相对频率,round(..., 4)是保留4位小数,你可以根据需求调整位数(比如改成2位)。 - 排序输出:用
sorted(...)按相对频率降序排列,和你原代码的输出顺序保持一致,方便对比查看。
示例输出
Processing 1 Relative frequencies: {'would': 0.2812, 'think': 0.1875, 'patient': 0.1719, 'condition': 0.1719, 'might': 0.1562, 'increased': 0.0156} Processing 2 Relative frequencies: {'cancer': 0.4156, 'condition': 0.2857, 'prostate': 0.2597, 'educational': 0.0129}
内容的提问来源于stack exchange,提问作者Shelina
相关产品推荐
相关产品推荐

