如何在考虑交易时段且无重绘的情况下将OHLC数据重采样为15分钟周期
解决Pandas重采样OHLC数据的交易时段与无重绘问题
问题描述
- 交易时段问题:使用
resample()时未严格遵循9:15 AM至3:30 PM的交易时段,闭市时段生成的NaN值导致Supertrend指标计算出错。 - 无重绘问题:未完成的15分钟周期(如当前时间为10:18 AM,对应10:15-10:30周期)不能使用未定型的聚合数据,需返回原始数据中上一条已完成的记录(如10:14 AM的3分钟数据)。
原代码
调用代码
prices_data_x_min_df = TechnicalHelper.resample_df(prices_data_df=prices_data_df, resample_period=self.supertrend_period_config) nifty_prices_data_x_min_df = TechnicalHelper.resample_df(prices_data_df=nifty_prices_data_df, resample_period=self.supertrend_period_config)
resample_df函数
@staticmethod def resample_df(prices_data_df: pd.DataFrame, resample_period: str): prices_data_df.set_index('record_ts', inplace=True) market_open = pd.Timestamp('9:15:00').time() market_close = pd.Timestamp('15:30:00').time() market_hours_data_df = prices_data_df.between_time(market_open, market_close) prices_data_x_min_df = market_hours_data_df.resample(resample_period).agg({ 'open': 'first', 'high': 'max', 'low': 'min', 'close': 'last', 'volume': 'sum' }, origin='start') prices_data_x_min_df.dropna(inplace=True) return prices_data_x_min_df
解决方案
核心思路
- 交易时段对齐:让重采样周期严格从每日9:15开始划分,仅保留交易时段内的完整周期,过滤闭市后无效时段。
- 无重绘处理:仅保留已完成的周期数据,对未完成的周期,直接复用原始数据的最后一条记录,避免使用未定型的聚合值。
修改后的resample_df函数
@staticmethod def resample_df(prices_data_df: pd.DataFrame, resample_period: str): # 复制数据避免修改原表 df = prices_data_df.copy() df.set_index('record_ts', inplace=True) market_open_time = pd.Timestamp('09:15:00').time() market_close_time = pd.Timestamp('15:30:00').time() # 过滤交易时段内的数据 market_hours_df = df.between_time(market_open_time, market_close_time) if market_hours_df.empty: return pd.DataFrame() last_data_ts = market_hours_df.index[-1] # 设置重采样起点为当日9:15,对齐交易时段 origin = pd.Timestamp(f"{last_data_ts.date()} 09:15:00") # 按指定周期重采样 resampled = market_hours_df.resample(resample_period, origin=origin).agg({ 'open': 'first', 'high': 'max', 'low': 'min', 'close': 'last', 'volume': 'sum' }) # 过滤闭市后的无效周期(如15:30之后的时段) resampled = resampled.between_time(market_open_time, market_close_time) # 标记并过滤未完成的周期 resampled['period_end'] = resampled.index + pd.Timedelta(resample_period) completed_periods = resampled[resampled['period_end'] <= last_data_ts].drop(columns='period_end') # 处理未完成周期:追加原始数据的最后一条记录 if last_data_ts < resampled['period_end'].iloc[-1]: last_raw_row = market_hours_df.iloc[-1].to_frame().T completed_periods = pd.concat([completed_periods, last_raw_row]) # 重置索引并恢复record_ts列 completed_periods.reset_index(inplace=True) completed_periods.rename(columns={'index': 'record_ts'}, inplace=True) return completed_periods
关键说明
- 交易时段严格过滤:通过
origin参数将重采样起点绑定到每日9:15,确保所有周期都落在交易时段内,同时用between_time二次过滤闭市后的无效周期。 - 无重绘保障:仅保留结束时间早于最后一条数据时间戳的已完成周期,未完成周期直接使用原始数据的最后一条记录,避免因聚合值后续变化导致的信号重绘。
- 安全操作:复制原数据进行处理,避免
inplace=True修改传入的原始数据集。
内容的提问来源于stack exchange,提问作者Shikhar Vaish
相关产品推荐
相关产品推荐

