如何为Pandas多级索引(MultiIndex)列的DataFrame添加计算列
优化实现方案
你可以通过列方向的groupby简化计算逻辑,代码更简洁,性能也更好:
import pandas as pd import numpy as np # 示例数据 arrays = [ np.array(["bar", "bar", "baz", "baz", "foo", "foo"]), np.array(["one", "two", "one", "two", "one", "two"]), ] df = pd.DataFrame(np.linspace(0, 4*6-1, 4*6).reshape(4,6), columns=arrays) # 按列0级索引分组,直接计算每个分组的one/two div_res = df.groupby(level=0, axis=1).apply( lambda x: x.xs('one', level=1, axis=1).values / x.xs('two', level=1, axis=1).values ) div_res.columns = pd.MultiIndex.from_product([div_res.columns, ['division']]) # 拼接数据 result = pd.concat([df, div_res], axis=1)
二级索引自定义排序
直接用reindex的level参数指定二级索引的顺序即可,一行完成:
# 按one → two → division的顺序排序二级索引 result = result.reindex(columns=['one', 'two', 'division'], level=1)
最终输出结果:
bar baz foo one two division one two division one two division 0 0 1 0.000000 2 3 0.666667 4 5 0.800000 1 6 7 0.857143 8 9 0.888889 10 11 0.909091 2 12 13 0.923077 14 15 0.933333 16 17 0.941176 3 18 19 0.947368 20 21 0.952381 22 23 0.956522
更直观的替代实现(适合小数据集)
如果你的数据量不大,还可以用stack把多级列索引转成行索引后直接计算,逻辑更易懂:
df_stacked = df.stack(level=0) df_stacked['division'] = df_stacked['one'] / df_stacked['two'] result = df_stacked.unstack().swaplevel(axis=1).sort_index(axis=1, level=0).reindex(columns=['one','two','division'], level=1)
内容的提问来源于stack exchange,提问作者divingTobi
相关产品推荐
相关产品推荐

