You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计DataFrame连续分组出现次数?已试cumcount/ngroup无效

解决连续分组计数问题

你需要的是连续相同值的分组编号,而非全局相同值的分组,直接用groupby的cumcount或ngroup确实满足不了需求,得先识别连续分组的边界:

步骤1:构造示例数据

import pandas as pd

df = pd.DataFrame({
    'id': [1,2,3,4,5,6,7,8],
    'webpage': ['google', 'bing', 'google', 'google', 'yahoo', 'yahoo', 'google', 'google']
})

步骤2:计算连续分组编号

通过判断当前行与上一行的webpage是否不同,生成布尔序列后累加,就能得到连续组的编号:

df['count'] = df['webpage'].ne(df['webpage'].shift()).cumsum()

最终结果

运行后得到的DataFrame完全符合你的需求:

id webpage  count
0   1   google      1
1   2    bing       2
2   3   google      3
3   4   google      3
4   5   yahoo       4
5   6   yahoo       4
6   7   google      5
7   8   google      5

原理说明:ne()是"not equal"的缩写,shift()将列数据向下偏移一行,webpage.ne(webpage.shift())会在每次webpage发生变化的行返回True(视为数值1),其余行返回False(视为数值0)。cumsum()对这个序列累加后,自然生成了连续分组的唯一编号。

内容的提问来源于stack exchange,提问作者AAk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 19:30:50