You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中分组两列取频次后获取最大值行

获取每个Topic下频次最高的Category

嘿,这需求很明确,咱们可以利用Pandas的分组和排序特性轻松解决。你现在通过df.groupby('topic')['category'].value_counts()得到的是一个带多层索引(topic + category)的Series,而且value_counts()默认已经按频次降序排列了,这刚好帮咱们省了排序的步骤~

方法一:直接取每组第一行(最简单)

因为value_counts()已经给每个topic下的category按频次从高到低排好序了,所以咱们只需要按外层的topic索引再分组,取每组的第一行就行:

# 先得到原始的频次统计
count_series = df.groupby('topic')['category'].value_counts()
# 按topic分组,取每组第一行(也就是频次最高的)
top_category = count_series.groupby(level='topic').head(1)

运行之后top_category就是你想要的结果:

topic   category   
topic1  Entertainment    1303
topic2  Politics         134
topic3  Entertainment    1370
dtype: int64

方法二:用idxmax精准定位(更严谨)

如果你担心默认排序的问题,或者需要明确指定取最大值的逻辑,可以用idxmax()找到每个topic下频次最高的那个索引组合,再提取对应行:

count_series = df.groupby('topic')['category'].value_counts()
# 找到每个topic下频次最大的索引((topic, category)元组)
max_indices = count_series.groupby(level='topic').idxmax()
# 根据索引提取对应的行
top_category = count_series.loc[max_indices]

这个方法的结果和上面完全一致,适合需要更明确逻辑的场景。

可选:转成规整的DataFrame格式

如果你希望结果是更易读的DataFrame(而非Series),可以加个.reset_index(name='count'):

top_category_df = top_category.reset_index(name='count')

得到的DataFrame结构如下:

topiccategorycount
topic1Entertainment1303
topic2Politics134
topic3Entertainment1370

内容的提问来源于stack exchange,提问作者Ronnie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:20:44