You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中定位分组最后一行并添加分组求和列?

问题描述

现有如下示例DataFrame:

firstsecondthirdfourthfifthsixth
aaa12332
aaa34613
bbb56945
bbb781217
ccc91015511
ccc111218913
ddd131421217

需要新增sum_of_sixth列,仅在每组(按first分组)的最后一行填入该组sixth列的求和值,其余行留空,预期结果如下:

firstsecondthirdfourthfifthsixthsum_of_sixth
aaa12332
aaa346135
bbb56945
bbb78121712
ccc91015511
ccc11121891324
ddd13142121717

尝试过concat和迭代方法未达到预期,求实现方案。

实现方案

方法一:分步实现

逻辑清晰,通过分组变换、标记最后行索引完成赋值:

  1. 用transform('sum')生成每组求和的序列,长度与原DataFrame一致
  2. 获取每组最后一行的索引
  3. 初始化新列为空,仅在最后行位置赋值求和值
import pandas as pd

# 构造示例数据
df = pd.DataFrame({
    'first': ['aaa', 'aaa', 'bbb', 'bbb', 'ccc', 'ccc', 'ddd'],
    'second': [1, 3, 5, 7, 9, 11, 13],
    'third': [2, 4, 6, 8, 10, 12, 14],
    'fourth': [3, 6, 9, 12, 15, 18, 21],
    'fifth': [3, 1, 4, 1, 5, 9, 2],
    'sixth': [2, 3, 5, 7, 11, 13, 17]
})

# 计算每组sixth列的求和值,每行对应所在组的总和
group_total = df.groupby('first')['sixth'].transform('sum')
# 获取每组最后一行的索引
last_row_indices = df.groupby('first').tail(1).index

# 新增列并赋值
df['sum_of_sixth'] = ''
df.loc[last_row_indices, 'sum_of_sixth'] = group_total.loc[last_row_indices]

print(df)

方法二:一行简化实现

利用cumcount(ascending=False)标记最后一行,结合where方法直接生成结果:

df['sum_of_sixth'] = df.groupby('first')['sixth'].transform('sum').where(
    df.groupby('first').cumcount(ascending=False) == 0,
    ''
)

cumcount(ascending=False)会给每组的行从末尾开始计数,最后一行的计数值为0,通过这个条件筛选出需要赋值的行,其余行设为空字符串。

内容的提问来源于stack exchange,提问作者Virgil Redfield

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 14:07:05