You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效按分组处理行实现pandas累计标记列,替代低效的iterrows方法

高效实现方案

你这个需求可以直接用pandas内置的分组累积最大值方法实现,完全不需要遍历行,性能远高于iterrows,百万级数据可以在毫秒级处理完成。

核心代码

import pandas as pd
import numpy as np

# 构造示例数据
df = pd.DataFrame({'student':'A A A B B B C C C'.split(),
                  'month':[1, 2, 3, 1, 2, 3, 1, 2, 3],
                  'pass':[0, 1, 0, 0, 0, 0, 1, 0, 0]})

# 直接生成pass_patch列
df['pass_patch'] = df.groupby('student')['pass'].cummax()

print(df)

输出结果

student  month  pass  pass_patch
0       A      1     0           0
1       A      2     1           1
2       A      3     0           1
3       B      1     0           0
4       B      2     0           0
5       B      3     0           0
6       C      1     1           1
7       C      2     0           1
8       C      3     0           1

和你预期的结果完全一致。

实现原理

cummax是累积最大值计算函数,配合groupby('student')按学生分组后,会在每个学生的分组内按行顺序计算pass列的累积最大值:

  • 只要该学生历史pass值出现过1,后续所有行的累积最大值都会保持为1
  • 如果从来没有出现过1,累积最大值就一直是0

完美匹配你的需求,所有运算都在pandas底层C语言层面执行,没有Python层面的循环,性能极高。

性能对比

对于百万级数据集:

  • iterrows逐行遍历需要数十秒到数分钟不等
  • 本方案只需要数十毫秒即可完成

内容的提问来源于stack exchange,提问作者Kaudrabear

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 03:09:03