Python如何统计列表中出现n次的单词数量
实现方法
不需要写复杂的自定义循环,不管是用标准库Counter还是Pandas,核心思路都是两层计数:先统计每个单词的出现次数,再统计「每个出现次数」对应的单词总数,两行核心代码就能搞定。
方案1:纯Python标准库实现(无第三方依赖,推荐)
用标准库collections.Counter即可,轻量高效,处理大列表性能很好:
from collections import Counter word_list = ["hello", "time", "burger", "hello", "mouse", "time", "time"] # 统计每个单词的出现频次 word_freq = Counter(word_list) # 统计每个频次对应的单词数量 freq_to_wordnum = Counter(word_freq.values()) # 按频次升序输出 for freq in sorted(freq_to_wordnum.keys()): count = freq_to_wordnum[freq] # 要输出整数就用下面这行,和你给的预期输出格式一致 print(f"{count} Word occurred {freq} time") # 如果需要你最开始提到的三位小数格式(比如4.500),替换成下面这行即可 # print(f"{count:.3f} Words appeared {freq} time")
运行后输出和你的预期完全匹配:
3 Word occurred 1 time 1 Word occurred 2 time 1 Word occurred 3 time
方案2:Pandas实现(适合已有数据处理流水线的场景)
如果你本身就在用Pandas处理文本数据,可以直接用Pandas自带的value_counts方法链式调用实现:
import pandas as pd word_list = ["hello", "time", "burger", "hello", "mouse", "time", "time"] # 链式做两次值计数,第一次得到单词->频次,第二次得到频次->单词数 freq_res = pd.Series(word_list).value_counts().value_counts().sort_index() for freq, word_count in freq_res.items(): print(f"{word_count} Word occurred {freq} time")
两种方案的时间复杂度都是O(n),处理十万级以上的单词列表都不会有性能问题,不需要额外写冗余的遍历逻辑。
内容的提问来源于stack exchange,提问作者Jakob QN
相关产品推荐
相关产品推荐

