You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中groupby结合resample时mean()与max()结果不一致的疑似Bug技术问询

Is this a Pandas Bug? Groupby + Resample Behavior Difference Between mean() and max()

First off, this isn't a bug — it's actually expected behavior in Pandas, rooted in how different aggregation functions handle non-numeric data. Let's break down what's happening:

Why the Two Results Differ

  • When using mean(): This aggregation function only operates on numeric columns. Your category column is string-type, so Pandas automatically excludes it from the calculation. The final result keeps category as part of the multi-index (from the initial groupby), with only the numeric value_hour column remaining as a data column.
  • When using max(): Unlike mean(), max() works with string columns (it returns the lexicographically largest value). Since each groupby('category') group only contains one unique category value, the max() of that group's category is just the value itself. This leads to category appearing both in the multi-index (from the groupby) and as a data column (from the max() aggregation on the string column).

Why drop('category') Throws a KeyError

The category you're trying to drop isn't a data column anymore — it's part of the DataFrame's multi-index. The drop() method by default searches for column names, so it can't locate 'category' in the columns, hence the KeyError.

Fixes to Match the mean() Output Style

You have a few straightforward options to get the result you want (with category only present in the index):

  1. Target the numeric column explicitly
    Call max() only on the column you care about, skipping the non-numeric category:

    df_max = df.groupby('category').resample('Y')['value_hour'].max()
    
  2. Use agg() for column-specific aggregation
    This gives you fine-grained control if you have multiple columns to aggregate:

    df_max = df.groupby('category').resample('Y').agg({'value_hour': 'max'})
    
  3. Clean up the existing multi-index DataFrame
    If you already have the duplicate category entry, you can adjust the index or reset and drop the redundant column:

    # Option 1: Remove the duplicate index level
    df_max_clean = df_max.droplevel('category')
    # Option 2: Reset index, drop the column, then re-set the index
    df_max_clean = df_max.reset_index().drop('category', axis=1).set_index(['category', 'level_1'])
    

Any of these approaches will produce a result consistent with your mean() output, where category only exists in the multi-index.

内容的提问来源于stack exchange,提问作者WhiteDear

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 14:09:19