You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中高效实现分组条件累加并按阈值重置?

问题描述

给定包含group、value、threshold列的DataFrame,需在每个group分组内对value列执行累加求和操作。每当累加和超过该分组指定的threshold阈值时,需重置累加状态,不同分组的threshold阈值可不同。

示例输入DataFrame:

import pandas as pd

data = {
    'group': ['A', 'A', 'A', 'A', 'B', 'B', 'B', 'C', 'C'],
    'value': [10, 20, 30, 40, 10, 20, 30, 40, 50],
    'threshold': [50, 50, 50, 50, 70, 70, 70, 100, 100]
}

df = pd.DataFrame(data)

对应的期望输出需新增cumsum_reset列:

group  value  threshold  cumsum_reset
0     A     10         50            10
1     A     20         50            30
2     A     30         50            30
3     A     40         50            40
4     B     10         70            10
5     B     20         70            30
6     B     30         70            60
7     C     40        100            40
8     C     50        100            90
高效实现方法

无需显式遍历每行,可通过Pandas分组功能结合itertools.accumulate实现高效计算,具体代码如下:

from itertools import accumulate

def reset_cumsum(group):
    # 定义累加逻辑:若累加后超过阈值则重置为当前value,否则继续累加
    def acc_logic(prev_sum, curr_pair):
        val, thresh = curr_pair
        new_sum = prev_sum + val
        return val if new_sum > thresh else new_sum
    
    # 对每组的(value, threshold)元组序列执行累加计算
    return list(accumulate(zip(group['value'], group['threshold']), acc_logic, initial=0))[1:]

# 分组计算并展开结果到原DataFrame
df['cumsum_reset'] = df.groupby('group').apply(reset_cumsum).explode().astype(int)

逻辑说明

  • itertools.accumulate是Python内置的高效累加工具,通过自定义的acc_logic函数判断是否需要重置累加状态,避免了Python层面的逐行循环
  • groupby.apply对每个分组独立处理,explode将分组生成的列表结果展开为对应行的数值,保证与原DataFrame的行对齐
  • 整体性能优于显式遍历,尤其适合处理大规模数据集

内容的提问来源于stack exchange,提问作者MarkusGreenberg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 03:11:03