Pandas使用groupby按词性分组统计高频词的实现方法
解决方法
你原来的写法返回的是多级索引的Series,没有转成你需要的字典结构,只需要加一步to_dict()转换即可,完整代码如下:
import pandas as pd # 你的示例数据 df = pd.DataFrame({ 'Parts of speech': ['Noun', 'Noun', 'Noun', 'verb', 'verb', 'adj'], 'word': ['cat', 'water', 'cat', 'draw', 'draw', 'slow'] }) # 实现统计 df2 = df.groupby('Parts of speech')['word']\ .apply(lambda x: x.value_counts().to_dict())\ .reset_index(name='top')
运行后输出df2就完全符合你的预期格式:
Parts of speech top 0 Noun {'cat': 2, 'water': 1} 1 adj {'slow': 1} 2 verb {'draw': 2}
如果不需要保留DataFrame结构,只需要得到词性对应词频字典的映射,也可以直接转成字典使用:
pos_count_dict = df.groupby('Parts of speech')['word'].apply(lambda x: x.value_counts().to_dict()).to_dict()
内容的提问来源于stack exchange,提问作者Rory
相关产品推荐
相关产品推荐

