如何对行列均为MultiIndex的DataFrame按指定列分组聚合?
解决MultiIndex行列DataFrame按指定列分组求和的问题
问题分析
你的DataFrame行、列均为MultiIndex,之前尝试的分组方法无效,是因为groupby默认操作行索引,且未明确指定列的MultiIndex完整路径。要实现按coords层级下的x、y分组,需要精准定位列的层级位置。
解决方案步骤
- 过滤有效数据行:排除
meta行(该行含字符串值,无求和意义),仅保留values组的数据。 - 指定分组键:用列的MultiIndex元组
("coords", "x")和("coords", "y")作为分组依据。 - 分组求和并整理结构:求和后重置索引,删除无用的
z列,调整列层级匹配目标格式。
完整代码
import pandas as pd # 构造原始DataFrame(你的代码) cols = pd.MultiIndex.from_arrays([*["coords"]*3, *["type"]*3],[ "x", "y", "z", "a", "b", "c"]) idx = pd.MultiIndex.from_arrays([*["values"]*6, "meta"], [*range(6), "foo"]) df = pd.DataFrame([[1,1,1,6,7,3], [1,1,0,1,5,9], [2,1,0,1,8,3], [2,1,0,5,7,2], [3,1,0,6,5,9], [3,1,0,7,4,5], [None, None, None, "bar", "baz", "qux"]], index=idx, columns=cols) # 1. 过滤values行 df_values = df.loc["values"] # 2. 按coords下的x、y分组求和 grouped_result = df_values.groupby([("coords", "x"), ("coords", "y")]).sum() # 3. 重置索引并调整列结构 result = grouped_result.reset_index().drop(("coords", "z"), axis=1) # 重新设置列的MultiIndex层级 result.columns = pd.MultiIndex.from_tuples([ ("coords", "x"), ("coords", "y"), ("type", "a"), ("type", "b"), ("type", "c") ]) print(result)
输出结果
coords type x y a b c 0 1.0 1.0 7 12 12 1 2.0 1.0 6 15 5 2 3.0 1.0 13 9 14
内容的提问来源于stack exchange,提问作者YPOC
相关产品推荐
相关产品推荐

