You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中pd.Grouper分组datetime列报错原因排查

Pandas按月份分组datetime列时出现TypeError错误

我在Python中使用Pandas,想要按月份对日期分组,并获取每个月dates_and_grades["Grade_Values"]列的最大值,编写的代码如下:

data = pd.read_csv(input_filepath)
data['Date'] = pd.to_datetime(data['Date'], format = 'ISO8601')

roped = ["Sport", "Trad"]

YDS_DICT={"N/A":"N/A",'3-4':0,'5':1,'5.0':1,'5.1':2,'5.2':3,'5.3':4,'5.4':5,
      '5.5':6,'5.6':7,'5.7':8,'5.8':9,'5.9':10,
      '5.10a':11,'5.10b':12, '5.10': 12, '5.10c':13,'5.10d':14,
      '5.11a':15,'5.11b':16, '5.11':16, '5.11c':17,'5.11d':18,
      '5.12a':19,'5.12b':20,'5.12c':21,'5.12d':22,
      '5.13a':23,'5.13b':24,'5.13c':25,'5.13d':26,
      '5.14a':27,'5.14b':28,'5.14c':29,'5.14d':30,
      '5.15a':31,'5.15b':32,'5.15c':33,'5.15d':34}

roped_only_naive = data.loc[data['Route Type'].isin(roped)].copy()
roped_only_naive["Rating"] = roped_only_naive['Rating'].map(slash_grade_converter)
roped_only_naive["Rating"] = roped_only_naive['Rating'].map(flatten_plus_and_minus_grades)
roped_only_naive["Rating"] = roped_only_naive['Rating'].map(remove_risk_ratings)
dates_and_grades = roped_only_naive[['Date', 'Rating']]
print(dates_and_grades.dtypes)
dates_and_grades["Grade_Values"] = dates_and_grades["Rating"].map(lambda data: YDS_DICT[data])
print(dates_and_grades.dtypes)
dates_and_grades['Date'] = dates_and_grades['Date'].groupby(pd.Grouper(freq='M'))
print(dates_and_grades)

运行后出现错误:

TypeError: Only valid with DatetimeIndex, TimedeltaIndex or PeriodIndex, but got an instance of 'Index'

通过print(dates_and_grades.dtypes)查看,Date列确实为datetime64[ns]类型:

Date            datetime64[ns]
Rating                  object
Grade_Values             int64

为何pd.Grouper(freq='M')分组操作在datetime类型的Date列上失效?


问题原因

出错的核心原因是:pd.Grouper(freq='M')默认是基于DataFrame的索引进行分组,而你的Date是普通列,不是索引。直接对Date列调用groupby(pd.Grouper(freq='M'))时,Pandas无法识别要按该列的时间维度分组,因此抛出错误。

正确解决方法

方法1:将Date列设为索引后分组

先把Date设置为DataFrame的索引,再使用pd.Grouper:

# 先设置Date为索引
dates_and_grades = dates_and_grades.set_index('Date')
# 按月份分组并取Grade_Values的最大值
monthly_max = dates_and_grades.groupby(pd.Grouper(freq='M'))['Grade_Values'].max()
print(monthly_max)

方法2:在groupby中指定key参数(无需修改索引)

直接在groupby里通过key参数指定要分组的Date列,同时使用pd.Grouper:

monthly_max = dates_and_grades.groupby(pd.Grouper(key='Date', freq='M'))['Grade_Values'].max()
print(monthly_max)

方法3:使用dt.to_period简化分组

更简洁的方式是利用datetime列的dt属性,直接转换为月份周期后分组:

monthly_max = dates_and_grades.groupby(dates_and_grades['Date'].dt.to_period('M'))['Grade_Values'].max()
print(monthly_max)

这三种方法都能实现按月份分组并获取Grade_Values的最大值,根据你的需求选择即可。

内容的提问来源于stack exchange,提问作者Luke

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 01:43:25