You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame多列分组取首项对应最大Date行方法

Pandas 多条件分组筛选行实现方案

原始数据

首先构造示例DataFrame:

import pandas as pd

df = pd.DataFrame(
    [[1,'a','c1',30,'s1','e1'],
     [1,'b','c1',60,'s1','e1'],
     [1,'b','c2',40,'s1','e1'],
     [2,'g','c1',40,'s2','e2'],
     [2,'g','c3',9,'s1','e1'],
     [3,'k','c2',20,'s1','e1'],
     [3,'k','c2',69,'s2','e1'],
     [3,'k','c1',29,'s1','e1'],
     [3,'f','c3',99,'s2','e1']],
    columns = ['Lot','Item','Code','Date','Shelf','Emp']
)

原始数据预览:

Lot Item Code  Date Shelf Emp
0    1    a   c1    30    s1  e1
1    1    b   c1    60    s1  e1
2    1    b   c2    40    s1  e1
3    2    g   c1    40    s2  e2
4    2    g   c3     9    s1  e1
5    3    k   c2    20    s1  e1
6    3    k   c2    69    s2  e1
7    3    k   c1    29    s1  e1
8    3    f   c3    99    s2  e1

需求拆解

需要依次完成3个操作:

  • 先按Lot字段、再按Item字段分组
  • 提取每个Lot分组下按原始顺序出现的首个Item值
  • 筛选每个Lot下对应首个Item的记录中,Date字段取最大值对应的整行数据

实现代码

采用Pandas原生的分组、映射、索引选择方法实现,避免低性能的逐行遍历:

# 1. 按Lot分组,取每组按原始顺序的第一个Item值,生成Lot到首个Item的映射
first_item_map = df.groupby('Lot', sort=False)['Item'].first()

# 2. 筛选出所有Item值等于对应Lot首个Item的记录
matched_subset = df[df['Item'] == df['Lot'].map(first_item_map)]

# 3. 对筛选后的子集按Lot分组,取Date最大值对应的行索引,提取整行后重置索引
result = matched_subset.loc[matched_subset.groupby('Lot')['Date'].idxmax()].reset_index(drop=True)

结果验证

运行代码后得到的结果如下:

Lot Item Code  Date Shelf Emp
0    1    a   c1    30    s1  e1
1    2    g   c1    40    s2  e2
2    3    k   c2    69    s2  e1

说明:问题描述中给出的预期输出Lot=2行的Code、Emp字段存在笔误,按照原始数据逻辑,Lot=2下首个Item为g,对应记录中Date最大值为40,匹配的Code为c1、Emp为e2,和代码运行结果一致。

内容的提问来源于stack exchange,提问作者AB Code

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.02 23:15:59