You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取pandas DataFrame分组后每组最后一行的全量字段数据

分组取最后一条全字段记录解决方案

问题背景

现有如下员工任职记录DataFrame:

PersID    Date   FirmID   Start   End   Match
     1    2017      100    2016  2017       1
     1    2017      200    2017  2019       1
     1    2018      200    2017  2019       1
     2    2016      600    2014  2017       1
     2    2017      600    2014  2017       1
     2    2017      700    2017  2017       1
     2    2017      800    2017  2019       1

需求为按PersID、Date两个字段分组,提取每组最后一条观测值,完整保留FirmID、Start、End、Match所有其余字段,期望输出如下:

PersID    Date   FirmID   Start   End   Match
     1    2017      200    2017  2019       1
     1    2018      200    2017  2019       1
     2    2016      600    2014  2017       1
     2    2017      800    2017  2019       1

之前使用的代码仅单独对Match列取聚合结果,因此丢失了其余字段:

df.groupby(['PersID', 'Date'])['Match'].last()

返回结果仅包含分组键和Match列,不符合要求:

PersID    Date   Match
     1    2017       1
     1    2018       1
     2    2016       1
     2    2017       1

实现方法

有两种常用写法可以实现全字段保留的分组取最后一条记录:

  • 方法1:直接对分组对象调用last()聚合,不要单独指定单列,添加as_index=False让分组键保持为普通列而非行索引
df_result = df.groupby(['PersID', 'Date'], as_index=False).last()
  • 方法2:使用tail()方法直接取每组最后N行,这里取1行即可,写法语义更直观
df_result = df.groupby(['PersID', 'Date'], as_index=False).tail(1)

两种方法运行后得到的结果和期望输出完全一致。

提示:如果需要保留原始DataFrame的行索引,去掉as_index=False参数即可。

内容的提问来源于stack exchange,提问作者acbcccdc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 06:54:23