如何在Pandas DataFrame的列中找出出现频率最高的词
问题描述
需要在Pandas DataFrame的Competition列中,为每行找出出现频率最高的词;如果有多个词出现次数相同且均为最高频,则将这些词全部用逗号分隔列出。
示例数据
import pandas as pd df = pd.DataFrame({ 'City': ['Pune', 'Mumbai', 'Pune', 'Mumbai', 'Pune'], 'Name': ['John', 'Boby', 'John', 'Boby', 'Nicky'], 'Competition': [ 'Chess,Drawing,Chess', 'Table Tennis,Table Tennis,Chess,Carrom', 'Chess,Carrom', 'Table Tennis,Chess,Chess,Chess', 'Carrom' ] })
原始数据:
| City | Name | Competition | |
|---|---|---|---|
| 0 | Pune | John | Chess,Drawing,Chess |
| 1 | Mumbai | Boby | Table Tennis,Table Tennis,Chess,Carrom |
| 2 | Pune | John | Chess,Carrom |
| 3 | Mumbai | Boby | Table Tennis,Chess,Chess,Chess |
| 4 | Pune | Nicky | Carrom |
期望输出
| City | Name | Competition | Most Frequent | |
|---|---|---|---|---|
| 0 | Pune | John | Chess,Drawing,Chess | Chess |
| 1 | Mumbai | Boby | Table Tennis,Table Tennis,Chess,Carrom | Table Tennis |
| 2 | Pune | John | Chess,Carrom | Carrom,Chess |
| 3 | Mumbai | Boby | Table Tennis,Chess,Chess,Chess | Chess |
| 4 | Pune | Nicky | Carrom | Carrom |
解决方案
通过自定义函数结合apply方法即可实现需求,代码如下:
import pandas as pd from collections import Counter def get_most_frequent(s): # 分割字符串为单个词的列表 items = s.split(',') # 统计每个词的出现次数 count_result = Counter(items) # 获取最高出现次数 max_freq = max(count_result.values()) # 筛选所有出现次数等于最高频的词 top_items = [word for word, freq in count_result.items() if freq == max_freq] # 用逗号连接结果返回 return ','.join(top_items) # 生成新列Most Frequent df['Most Frequent'] = df['Competition'].apply(get_most_frequent) # 查看结果 print(df)
代码说明
- 分割字符串:用
split(',')把每行的Competition内容拆分成独立的词列表; - 统计频次:借助
Counter快速统计每个词的出现次数; - 筛选最高频词:先找到最大出现次数,再筛选出所有达到该次数的词;
- 生成结果列:通过
apply将函数作用于Competition列的每一行,生成目标列。
运行上述代码后,就能得到符合要求的DataFrame。
内容的提问来源于stack exchange,提问作者anjali joshi
相关产品推荐
相关产品推荐

