Pandas按条件分组实现列cumulative sum累积求和方法
Pandas分组逐行累计求和实现
直接使用pandas的分组累计求和接口即可完成需求,以下是可直接运行的完整代码:
import pandas as pd # 构建示例数据集 df = pd.DataFrame({ 'innings': [1, 1, 1, 2, 2], 'matchid': [1000887, 1000887, 1000887, 1000887, 1000887], 'over_count': [1, 2, 50, 1, 50], 'runs_total_inover': [1.0, 1.0, 3.0, 6.0, 2.0] }) # 按innings、matchid分组,计算runs_total_inover的逐行累计和存入sum列 df['sum'] = df.groupby(['innings', 'matchid'])['runs_total_inover'].cumsum() print(df)
运行代码后输出结果和预期完全一致:
innings matchid over_count runs_total_inover sum 0 1 1000887 1 1.0 1.0 1 1 1000887 2 1.0 2.0 2 1 1000887 50 3.0 5.0 3 2 1000887 1 6.0 6.0 4 2 1000887 50 2.0 8.0
补充说明:
cumsum()计算时会遵循DataFrame当前的行顺序,如果你的原始数据行顺序不是按over_count升序排列的,建议在计算前先执行排序,避免累计结果出错:# 按分组维度、轮次排序后再计算累计和 df = df.sort_values(by=['innings', 'matchid', 'over_count'], ignore_index=True) df['sum'] = df.groupby(['innings', 'matchid'])['runs_total_inover'].cumsum()
内容的提问来源于stack exchange,提问作者KHURRAM
相关产品推荐
相关产品推荐

