You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas时间序列聚合如何设置分组起始点为第一行时间

解决方法

你只需要在pd.Grouper中新增offset参数,将时间分组窗口整体偏移5分钟即可,调整后的代码如下:

import pandas as pd

tf = 10
df_f = df.groupby(['script_id', pd.Grouper(key='date_time', freq=f'{tf}T', offset='5T')])\
                            .agg(open=pd.NamedAgg(column='open', aggfunc='first'),
                                high=pd.NamedAgg(column='high', aggfunc='max'),
                                low=pd.NamedAgg(column='low', aggfunc='min'),
                                close=pd.NamedAgg(column='close', aggfunc='last'),
                                volume=pd.NamedAgg(column='volume', aggfunc='sum'))\
                                .reset_index()
print(df_f)

默认情况下pandas的时间分组会对齐到频率的整数倍起始点,10分钟频率默认就对齐到0、10、20…分钟,加5分钟偏移后就对齐到5、15、25…分钟,刚好符合你的需求。

如果你使用的是pandas 1.1.0之前的旧版本,不支持offset参数,也可以通过修改时间列后再分组的方式实现:

df['adjusted_time'] = df['date_time'] - pd.Timedelta(minutes=5)
df_f = df.groupby(['script_id', pd.Grouper(key='adjusted_time', freq=f'{tf}T')])\
                            .agg(open=pd.NamedAgg(column='open', aggfunc='first'),
                                high=pd.NamedAgg(column='high', aggfunc='max'),
                                low=pd.NamedAgg(column='low', aggfunc='min'),
                                close=pd.NamedAgg(column='close', aggfunc='last'),
                                volume=pd.NamedAgg(column='volume', aggfunc='sum'))\
                                .reset_index()
# 回改回目标分组起始时间
df_f['date_time'] = df_f['adjusted_time'] + pd.Timedelta(minutes=5)
df_f = df_f.drop('adjusted_time', axis=1)

内容的提问来源于stack exchange,提问作者Kabomi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 12:00:01