You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas DataFrame中按sample_type统计连续值1的最大出现次数

按分组统计连续值1的最大出现次数

原始数据

import pandas as pd

df = pd.DataFrame({
    'sample_type': ['A', 'A', 'A', 'B', 'B', 'B', 'B', 'C', 'C', 'C'],
    'data_window': [1, 1, 1, 1, 1, 8, 1, 2, 1, 1]
})

实现方案

核心思路是先标记每组内连续1的分组,再统计各分组长度,最后取最大值:

# 1. 按sample_type分组,生成连续1的分组标识
df['group_key'] = df.groupby('sample_type')['data_window'].apply(
    lambda x: (x != 1 | x.shift() != x).cumsum()
)

# 2. 过滤出值为1的行,统计每个连续分组的长度
continuous_counts = df[df['data_window'] == 1].groupby(['sample_type', 'group_key']).size()

# 3. 按sample_type取最大连续长度,整理成目标格式
result = continuous_counts.groupby('sample_type').max().reset_index(name='cum_count')

代码解释

  • 生成group_key:在每个sample_type分组内,当data_window不等于1,或者当前值与前一行值不同时,生成新的分组编号,确保连续的1被归为同一个分组。
  • 统计连续长度:过滤出所有值为1的行,按sample_type和group_key统计行数,得到每组连续1的长度。
  • 取最大值:按sample_type分组后取最大的连续长度,用reset_index转换为目标DataFrame格式。

输出结果

sample_type  cum_count
0           A          3
1           B          2
2           C          2

内容的提问来源于stack exchange,提问作者MAtennis9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 22:12:23