Pandas新增列计算多级表头下所有one、two列的行列总和
Pandas多级列索引按二级维度统计行、列总和实现
实现步骤
- 首先构造和示例一致的多级列DataFrame
- 按二级列名分组,逐行计算所有
one列、two列的总和,作为新列追加 - 追加Total汇总行,逐列计算所有数值列的总和
完整代码
import pandas as pd # 1. 构造源数据 data = [ [0.895717, 0.805244, 1.206412, 2.565646, 1.431256, 1.431256], [0.410835, 0.813850, 0.132003, 0.827317, 0.076467, 0.076467], [1.413681, 1.607920, 1.024180, 0.569605, 0.875906, 0.875906] ] # 构建二级列索引 columns = pd.MultiIndex.from_product([['bar', 'baz', 'foo'], ['one', 'two']]) df = pd.DataFrame(data, index=['A', 'B', 'C'], columns=columns) df.index.name = 'name' # 2. 新增行维度统计列:逐行统计所有二级列为one、two的总和 df[('', 'total of one')] = df.xs('one', axis=1, level=1).sum(axis=1) df[('', 'total of two')] = df.xs('two', axis=1, level=1).sum(axis=1) # 3. 新增Total汇总行:逐列统计总和 df.loc['Total'] = df.sum(numeric_only=True, axis=0)
运行结果
bar baz foo one two one two one two total of one total of two name A 0.895717 0.805244 1.206412 2.565646 1.431256 1.431256 3.533385 4.802146 B 0.410835 0.813850 0.132003 0.827317 0.076467 0.076467 0.619305 1.717634 C 1.413681 1.607920 1.024180 0.569605 0.875906 0.875906 3.313767 3.053431 Total 2.720233 3.227014 2.362595 3.962568 2.383629 2.383629 7.466457 9.573211
说明
- 用*
xs切片方法*处理多级列索引,不需要手动枚举一级列名,后续新增一级列时统计逻辑不需要修改,扩展性更好 - 新增统计列时一级列名传空值,可完美适配原有二级列结构,不会破坏列索引格式
- 汇总行计算时加
numeric_only=True参数,可自动跳过非数值列,避免运行报错
内容的提问来源于stack exchange,提问作者ZAVERI SIR
相关产品推荐
相关产品推荐

