You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas:保留多层索引中第二层索引重复项的首行

按多层索引的第二层ID去重并保留首行数据

需求说明

针对带有多层索引(date为第一层,ID为第二层)的DataFrame,当第二层索引ID存在重复项时,保留每个唯一ID对应的首次出现行,忽略第一层索引date的差异。

示例数据

你的示例数据结构如下:

col_0col_1col_2col_3col_4
dateID
('2022-01-01', 'identifier_0')2646442110
('2022-01-01', 'identifier_1')2545832345
('2022-01-01', 'identifier_2')427955578
('2022-01-01', 'identifier_3')324571961
('2022-01-01', 'identifier_4')302559372
('2022-01-02', 'identifier_0')4214564342
('2022-01-02', 'identifier_1')902746585
('2022-01-02', 'identifier_2')3339539486
('2022-01-02', 'identifier_3')3265988164
('2022-01-02', 'identifier_4')4831255815
('2022-01-03', 'identifier_0')580339680
('2022-01-03', 'identifier_1')1586453962
('2022-01-03', 'identifier_2')983425083

解决方案

假设你的DataFrame变量名为df,多层索引的第二层名称为ID,直接使用drop_duplicates方法即可实现需求:

# 按ID去重,保留每个ID首次出现的行
df_unique = df.drop_duplicates(subset='ID', keep='first')

参数说明

  • subset='ID':指定以第二层索引ID作为重复判断的依据
  • keep='first':保留每个重复ID对应的第一行,即该ID在数据中首次出现的记录(对应最早的date)

处理后结果

执行上述代码后,得到的df_unique将包含以下行(每个ID仅保留首次出现的记录):

col_0col_1col_2col_3col_4
('2022-01-01', 'identifier_0')2646442110
('2022-01-01', 'identifier_1')2545832345
('2022-01-01', 'identifier_2')427955578
('2022-01-01', 'identifier_3')324571961
('2022-01-01', 'identifier_4')302559372

内容的提问来源于stack exchange,提问作者Zen4ttitude

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 14:05:30