You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

聚合函数生成的DataFrame转换:拆分amin/amax为独立列

Pandas 多层列转换为扁平化独立列

问题场景

通过以下代码生成了带有多层列索引的DataFrame:

import numpy as np
import pandas as pd

# 假设df为原始数据集
aa = df.groupby(['ID_number']).agg({'PushedDate': [np.min, np.max]})
aa[aa.index.name] = aa.index 
aa.index.names = ['index']  # 重命名索引

生成的DataFrame结构如下:

PushedDate                ID_number
                   amin         amax    
index           
7874            2022-02-18  2022-02-22       7874
4979            2021-10-06  2021-10-11       4979

需要将其转换为扁平化列结构,让amin和amax作为独立列,最终效果如下:

PushedDate_min   PushedDate_max   ID_number  
index           
7874            2022-02-18        2022-02-22         7874
4979            2021-10-06        2021-10-11         4979

解决方案

方法1:聚合阶段直接指定目标列名(推荐)

从根源避免生成多层列,在agg中直接定义每个聚合操作对应的列名:

# 分组聚合时直接设置目标列名,无需后续处理多层列
aa = df.groupby('ID_number', as_index=False).agg(
    PushedDate_min=('PushedDate', np.min),
    PushedDate_max=('PushedDate', np.max)
)
# 按需求设置索引名称
aa = aa.set_index('ID_number').rename_axis('index')

执行后直接得到目标结构,步骤更简洁。

方法2:对已生成的多层列DataFrame做扁平化处理

如果已经有了用户提供的aa,可以通过拆分多层列并重新命名来实现:

# 拆分PushedDate的多层列,并重命名为目标列名
pushed_flat = aa['PushedDate'].rename(columns={
    'amin': 'PushedDate_min',
    'amax': 'PushedDate_max'
})
# 合并扁平化后的列与ID_number列
result = pd.concat([pushed_flat, aa['ID_number']], axis=1)

方法3:批量拼接多层列名

通过列表推导式批量拼接多层列的名称,适合通用场景:

# 将多层列的两级名称用下划线拼接,单层列直接保留原名称
aa.columns = ['_'.join(col).strip() for col in aa.columns.values]
# 可选:将PushedDate_amin/amax重命名为更贴合需求的名称
aa = aa.rename(columns={
    'PushedDate_amin': 'PushedDate_min',
    'PushedDate_amax': 'PushedDate_max'
})

内容的提问来源于stack exchange,提问作者Roshankumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 07:15:51