如何简洁扁平化Pandas Groupby聚合后的多级索引列?
Great question! Manually renaming columns after groupby-aggregation is tedious and error-prone, especially as your aggregation logic grows. Let's cover two straightforward approaches to avoid or flatten those multi-level columns automatically:
1. Define Column Names Directly in agg() (Best Practice)
Instead of letting Pandas generate multi-level columns, you can specify custom column names while aggregating using tuple syntax. This skips the multi-level index entirely:
import pandas as pd import numpy as np df = pd.DataFrame( {'A': [1,1,1,2,2,2,3,3,3], 'B': np.random.random(9), 'C': np.random.random(9)} ) # Specify column names alongside aggregation functions out = df.groupby('A').agg( B_mean=('B', np.mean), B_std=('B', np.std), C_median=('C', np.median) ) print(out)
This gives you the flattened columns directly, no post-processing needed:
B_mean B_std C_median A 1 0.791846 0.091657 0.394167 2 0.156290 0.202142 0.453871 3 0.482282 0.382391 0.892514
2. Flatten Existing Multi-Level Columns
If you already have a DataFrame with multi-level columns (like your original example), you can automate renaming by joining the level values:
Option A: List Comprehension with str.join
# Using your original out DataFrame with multi-level columns out.columns = ['_'.join(col) for col in out.columns]
Option B: Use Pandas' get_level_values
This is useful if you want more control over how levels are combined:
out.columns = out.columns.get_level_values(0) + '_' + out.columns.get_level_values(1)
Both methods will produce the same flattened column names as your manual approach, but without the hassle of typing each name manually.
Bonus: Save to Text File
Once your columns are flattened, saving to a text file (like CSV or tab-separated) is straightforward:
out.to_csv('aggregated_data.txt', sep='\t') # Use sep=',' for standard CSV format
内容的提问来源于stack exchange,提问作者Haleemur Ali

