Pandas中groupby结合resample时mean()与max()结果不一致的疑似Bug技术问询
mean() and max() First off, this isn't a bug — it's actually expected behavior in Pandas, rooted in how different aggregation functions handle non-numeric data. Let's break down what's happening:
Why the Two Results Differ
- When using
mean(): This aggregation function only operates on numeric columns. Yourcategorycolumn is string-type, so Pandas automatically excludes it from the calculation. The final result keepscategoryas part of the multi-index (from the initialgroupby), with only the numericvalue_hourcolumn remaining as a data column. - When using
max(): Unlikemean(),max()works with string columns (it returns the lexicographically largest value). Since eachgroupby('category')group only contains one uniquecategoryvalue, themax()of that group'scategoryis just the value itself. This leads tocategoryappearing both in the multi-index (from thegroupby) and as a data column (from themax()aggregation on the string column).
Why drop('category') Throws a KeyError
The category you're trying to drop isn't a data column anymore — it's part of the DataFrame's multi-index. The drop() method by default searches for column names, so it can't locate 'category' in the columns, hence the KeyError.
Fixes to Match the mean() Output Style
You have a few straightforward options to get the result you want (with category only present in the index):
Target the numeric column explicitly
Callmax()only on the column you care about, skipping the non-numericcategory:df_max = df.groupby('category').resample('Y')['value_hour'].max()Use
agg()for column-specific aggregation
This gives you fine-grained control if you have multiple columns to aggregate:df_max = df.groupby('category').resample('Y').agg({'value_hour': 'max'})Clean up the existing multi-index DataFrame
If you already have the duplicatecategoryentry, you can adjust the index or reset and drop the redundant column:# Option 1: Remove the duplicate index level df_max_clean = df_max.droplevel('category') # Option 2: Reset index, drop the column, then re-set the index df_max_clean = df_max.reset_index().drop('category', axis=1).set_index(['category', 'level_1'])
Any of these approaches will produce a result consistent with your mean() output, where category only exists in the multi-index.
内容的提问来源于stack exchange,提问作者WhiteDear

