You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas resample规则解析及30分钟滚动窗口重采样实现疑问

Pandas resample规则解析及30分钟滚动窗口重采样实现疑问

你观察得太到位了!这其实是Pandas resample() 方法的核心规则导致的——它默认是锚定标准时间刻度的固定频率分箱,而非你需要的滚动滑动窗口,下面给你拆解清楚:

一、先搞懂resample的分箱规则

1. 标准频率的锚定行为

当你使用像30min、1h这种和系统时间刻度对齐的频率时,resample会严格按照全球标准时间刻度来划分时间箱:

  • 30分钟频率的分箱锚点是00:00、00:30、01:00、01:30……以此类推,不管你的数据起始/结束时间是什么,所有分箱都会对齐到这些点。
  • 这就解释了为什么当前时间是9:28时,最后一个分箱是9:00-9:30,但因为这个时间箱还没到截止点(9:30),所以resample不会把9:00到9:28的数据单独生成一个未完成的分箱,只会输出到上一个完整的标准刻度分箱。

2. 非标准频率的递推行为

而当你用29min这种“非标准”频率时,Pandas找不到对应的全局锚点,就会从你的数据集的起始时间开始,以指定频率向后递推分箱,比如第一个分箱是起始时间~起始时间+29min,第二个是起始时间+29min~起始时间+58min……看起来就像是滚动的,但本质上还是固定步长的分箱,并非真正的“滑动窗口”(滑动窗口是每个时间点都对应一个窗口,而这个是每29分钟一个固定块)。

你给出的两段示例代码完美验证了这个差异,整理后如下:

示例1:随机时间序列的对比

import pandas as pd
import numpy as np

recent_hours = 2
# 生成间隔3.4分钟的时间序列,截止到当前时间
date_index = pd.date_range(end=pd.to_datetime('now'), periods=recent_hours * 60 // 3.4, freq='3T')

# 构造模拟OHLC数据
data = pd.DataFrame({
    'open': np.random.rand(len(date_index)),
    'high': np.random.rand(len(date_index)),
    'low': np.random.rand(len(date_index)),
    'close': np.random.rand(len(date_index))
}, index=date_index)

# 30分钟重采样:严格对齐整点/半点
sample_30m = data.resample('30min').agg({
    'open': 'first',
    'high': 'max',
    'low': 'min',
    'close': 'last',
})

# 29分钟重采样:从数据起始时间递推分箱
sample_29m = data.resample('29min').agg({
    'open': 'first',
    'high': 'max',
    'low': 'min',
    'close': 'last',
})

示例2:固定时间点的对比

import pandas as pd
import datetime

# 构造带固定时间索引的OHLC数据
df= [[192, 203, 188, 193],
[195, 201, 178, 196],
[196, 210, 182, 201],
[199, 211, 198, 203],
[201, 221, 197, 212],
[203, 218, 195, 209]]

data = pd.DataFrame(df, columns=['open', 'high', 'low', 'close'],
index=[
datetime.datetime(2024, 1, 11, 8, 57, 32),
datetime.datetime(2024, 1, 11, 9, 13, 12),
datetime.datetime(2024, 1, 11, 9, 25, 8),
datetime.datetime(2024, 1, 11, 9, 32, 56),
datetime.datetime(2024, 1, 11, 9, 42, 32),
datetime.datetime(2024, 1, 11, 10, 2, 11)
])

sample_30m = data.resample('30min').agg({
    'open': 'first',
    'high': 'max',
    'low': 'min',
    'close': 'last',
})

sample_29m = data.resample('29min').agg({
    'open': 'first',
    'high': 'max',
    'low': 'min',
    'close': 'last',
})

运行后你会明显看到:sample_30m的索引全是整点/半点,而sample_29m的索引从第一个数据点的时间开始递推,完全不跟标准时间对齐。

二、实现真正的30分钟滚动窗口重采样

如果你想要的是包含最新数据的滑动窗口(比如当前时间是9:28,窗口就是8:58-9:28),而非固定刻度的分箱,那resample()就不太适用了,应该用rolling()方法来实现:

方法:用rolling()做滑动窗口聚合

# 30分钟滚动窗口,按时间定义窗口而非数据点数
rolling_30m = data.rolling('30min', closed='right').agg({
    'open': lambda x: x.iloc[0] if len(x) > 0 else np.nan,
    'high': 'max',
    'low': 'min',
    'close': lambda x: x.iloc[-1] if len(x) > 0 else np.nan,
})

# 如果你想要类似resample的稀疏输出(而非每个数据点都有窗口结果),可以结合resample和rolling
# 比如每5分钟取一次30分钟窗口的聚合结果
resampled_rolling = data.resample('5min').agg({
    'open': lambda x: data.loc[x.index[0]-pd.Timedelta('30min'):x.index[-1], 'open'].iloc[0],
    'high': lambda x: data.loc[x.index[0]-pd.Timedelta('30min'):x.index[-1], 'high'].max(),
    'low': lambda x: data.loc[x.index[0]-pd.Timedelta('30min'):x.index[-1], 'low'].min(),
    'close': lambda x: data.loc[x.index[0]-pd.Timedelta('30min'):x.index[-1], 'close'].iloc[-1],
})

这里要注意:

  • rolling('30min')是真正的滑动窗口,每个时间点都会对应一个向前30分钟的窗口
  • 用lambda函数模拟resample的OHLC聚合逻辑,取窗口第一个open和最后一个close
  • closed='right'参数控制窗口包含右边界(即当前时间点属于窗口内)

三、补充:用resample实现类滚动效果

如果你还是想用resample(),可以通过origin参数修改分箱锚点,比如把锚点设为数据的最后一个时间点,让分箱从最后时间往前推:

# 把resample的锚点设为数据的最后一个时间点
sample_30m_rolling = data.resample('30min', origin='end').agg({
    'open': 'first',
    'high': 'max',
    'low': 'min',
    'close': 'last',
})

这样生成的分箱会以最后一条数据的时间为终点向前递推,看起来类似滚动,但本质还是固定步长的分箱,和rolling的滑动窗口有本质区别。

备注:内容来源于stack exchange,提问作者Chen Sullivan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.20 06:18:03