如何基于Pandas DataFrame统计每个用户的重复评论数量?
高效统计Pandas DataFrame中用户的重复评论
嗨,这事儿用Pandas的原生分组聚合就能高效搞定,速度快还省心!我给你一步步拆解:
首先,先把你的原始数据转成可直接使用的DataFrame结构:
import pandas as pd data = { 'Author': ['casy', 'linda', 'casy', 'tom', 'bob', 'bob', 'bob', 'bob', 'casy', 'casy', 'linda', 'linda', 'bob', 'casy'], 'Comment': ['Nice picture!', 'I like this', 'Nice picture!', 'I disagree', 'Follow me', 'Follow me', 'Follow me', 'Follow me', 'Nice picture!', 'Wow!', 'Interesting post', 'Check my profile', 'Dissapointing', 'Wow!'] } df = pd.DataFrame(data)
接下来是核心操作,一行链式代码直接得到你要的结果:
result = (df.groupby(['Author', 'Comment'], as_index=False) .size() .rename(columns={'size': 'Number of dup. comments (descending)'}) .sort_values(by=['Author', 'Number of dup. comments (descending)'], ascending=[True, False]))
我给你解释下每一步的作用:
groupby(['Author', 'Comment']):把同一个用户+同一条评论的所有行归为一组,这是统计重复评论的核心逻辑.size():直接统计每组的行数,也就是这条评论的重复次数rename(...):把默认的统计列名size改成你需要的自定义列名sort_values(...):先按用户名升序排列,再按重复次数降序排列,让同一个用户的评论按重复频次从高到低展示,完全匹配你的需求
如果想要输出和你示例一致的纯文本表格样式,用to_string(index=False)即可:
print(result.to_string(index=False))
最终输出结果如下:
Author Number of dup. comments (descending) Comment bob 4 Follow me bob 1 Dissapointing casy 3 Nice picture! casy 2 Wow! linda 1 I like this linda 1 Check my profile linda 1 Interesting post tom 1 I disagree
这个方案完全依赖Pandas原生优化函数,没有冗余循环,大数据量下也能保持高效运行,绝对是最快的解决方案之一。
内容的提问来源于stack exchange,提问作者user40
相关产品推荐
相关产品推荐

