You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何避免遍历多级索引DataFrame的level=0列名添加求和列?

多级索引DataFrame批量添加层级内求和列

先看初始的DataFrame结构:

import pandas as pd
import numpy as np

# 创建带多级列索引的DataFrame
df = pd.DataFrame(np.random.randint(0,10,[5,4]), 
                   columns=pd.MultiIndex.from_product([['A','B'], ['a','b']]))
print(df)

输出结果:

A       B
        a   b   a   b
0   6   0   7   3
1   0   5   1   8
2   3   4   6   4
3   6   2   0   2
4   3   9   5   0

需求是给每个level=0的分组(比如A、B)新增一个level=1为c的列,值为该分组下所有level=1列的行求和。如果手动逐个指定level=0列名,写法如下:

# 手动给A、B分组添加求和列
df.loc[:, pd.IndexSlice['A','c']] = df.loc[:, pd.IndexSlice['A',:]].sum(axis=1)
df.loc[:, pd.IndexSlice['B','c']] = df.loc[:, pd.IndexSlice['B',:]].sum(axis=1)
df = df.sort_index(axis=1)
print(df)

输出:

A           B
        a   b   c   a   b   c
0   6   0   6   7   3   10
1   0   5   5   1   8   9
2   3   4   7   6   4   10
3   6   2   8   0   2   2
4   3   9   12  5   0   5

但实际场景中level=0有数百个列名,手动遍历效率太低,以下是无需遍历的批量实现方法(适配Python3.8、pandas1.4.4):


方法一:groupby求和+concat合并

核心思路是按level=0分组求和,再把结果转换成对应多级索引格式,最后和原数据合并排序:

# 1. 按level=0分组,计算每行的求和值
sum_result = df.groupby(level=0, axis=1).sum()
# 2. 给求和结果的列构造多级索引,level=1设为'c'
sum_result.columns = pd.MultiIndex.from_tuples([(col, 'c') for col in sum_result.columns])
# 3. 合并原DataFrame和求和列,再按列索引排序
df = pd.concat([df, sum_result], axis=1).sort_index(axis=1)

方法二:groupby+rename简化写法

利用rename直接修改列的多级索引,写法更简洁:

# 分组求和后,直接把列名重命名为(原level0值, 'c')
sum_cols = df.groupby(level=0, axis=1).sum().rename(columns=lambda x: (x, 'c'))
# 合并并排序索引
df = pd.concat([df, sum_cols], axis=1).sort_index(axis=1)

这两种方法都能自动处理所有level=0的分组,不需要手动遍历列名,效率更高,适合level=0列数较多的场景。


内容的提问来源于stack exchange,提问作者data-monkey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 16:13:19