You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas如何对带时区的datetime列按小时分组求和

问题原因

你使用pd.Grouper(key='StartDate', freq='H')分组时报错,核心原因是StartDate列同时存在-04:00、-02:00两种不同的时区偏移量,pandas无法在混合时区的场景下直接按小时频率做自动分箱,会将该列识别为普通Index类型而非可重采样的DatetimeIndex,最终抛出类型错误:

TypeError: Only valid with DatetimeIndex, TimedeltaIndex or PeriodIndex, but got an instance of 'Index'

解决方案

如果要实现你预期的分组效果(按每条记录的本地小时分组,同个本地小时的记录归为一组、Weight求和,最终输出3行结果),不需要依赖Grouper,直接对时间列做小时维度向下取整作为分组键即可,代码如下:

import pandas as pd

df = pd.DataFrame(
    {
        'StartDate':['2020-01-01 00:00:00-04:00', '2020-01-01 01:00:00-04:00', '2020-01-01 01:55:00-04:00', '2020-01-02 02:00:00-02:00', '2020-01-02 02:00:00-04:00'],
        'Weight':[100, 110, 120, 125, 155]
    }
)
df['StartDate'] = pd.to_datetime(df['StartDate'])

# 按时间的本地小时向下取整后分组求和
result = df.groupby(df['StartDate'].dt.floor('H'))['Weight'].sum()

运行后输出结果完全匹配预期:

StartDate
2020-01-01 00:00:00-04:00    100
2020-01-01 01:00:00-04:00    230
2020-01-02 02:00:00-02:00    280
Name: Weight, dtype: int64

如果你一定要使用pd.Grouper实现分组,必须先将所有时间统一转换到同一个时区(比如UTC时区),再传入Grouper即可正常运行,示例代码:

# 统一转换为UTC时区
df['StartDate'] = df['StartDate'].dt.tz_convert('UTC')
result = df.groupby(pd.Grouper(key='StartDate', freq='H'))['Weight'].sum()

注意:统一时区后是按时区转换后的标准时间小时分组,分组结果和你预期的本地小时分组结果存在差异,根据实际业务场景选择即可。


内容的提问来源于stack exchange,提问作者IamWarmduscher

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 02:36:08