如何对带MultiIndex列的pandas DataFrame按指定层级分组
解决MultiIndex列的DataFrame按指定层级列分组的问题
当DataFrame的列采用MultiIndex结构,想要按某一层级(比如baz)下的所有列的值分组时,直接用groupby("baz")会报错——因为groupby默认会把传入的字符串当作行索引的层级,而非列的层级。下面给出两种实用的解决方案:
方法一:动态筛选目标层级列(灵活通用)
先筛选出first层级为baz的所有列,再将这些列作为分组依据传入groupby,适合列结构可能变动的场景:
import pandas as pd # 构造示例数据 iterables = [["bar", "baz", "foo", "qux"], ["one", "two"]] idx = pd.MultiIndex.from_product(iterables, names=["first", "second"]) data = [ [0, 0, 1, 1, 2, 3, 0, 0], [0, 0, 1, 1, 3, 2, 0, 0], [0, 1, 2, 3, 1, 2, 0, 0], [1, 0, 2, 3, 0, 3, 0, 0], ] df_with_multi = pd.DataFrame(data=data, columns=idx) # 筛选first层级为baz的列 baz_group_cols = df_with_multi.columns[df_with_multi.columns.get_level_values("first") == "baz"] # 按这些列分组,取每组第一条数据 result = df_with_multi.groupby(baz_group_cols.tolist()).first() # 调整索引名称,匹配期望输出的显示格式 result.index.names = ["one", "two"] result.index = result.index.rename("baz", level=0) print(result)
输出结果:
first bar foo qux second one two one two one two baz one two 1 1 0 0 2 3 0 0 2 3 0 1 1 2 0 0
方法二:直接指定列的元组标签(适合已知列名的场景)
如果你明确知道baz下的子层级列名,可以直接传入元组形式的列标签,写法更简洁:
result = df_with_multi.groupby([("baz", "one"), ("baz", "two")]).first() # 调整索引名称以匹配期望格式 result.index.names = ["one", "two"] result.index = result.index.rename("baz", level=0)
报错原因说明
df_with_multi.groupby("baz").first()报错ValueError: Grouper for 'baz' not 1-dimensional,核心原因是:groupby默认将字符串参数解析为行索引的层级名称,而你的baz是列MultiIndex的层级,并非行索引层级,因此无法识别。必须明确指定要分组的列(无论是动态筛选还是直接指定列标签),才能实现按目标层级列分组的需求。
内容的提问来源于stack exchange,提问作者spettekaka
相关产品推荐
相关产品推荐

