如何避免Pandas Resample生成额外小时数据?仅保留原时间范围数据
问题
我用Pandas的resample函数将分钟级数据转换为小时级数据,原始DataFrame的时间范围仅为10:30至15:59,但重采样后生成了额外的小时数据,希望移除这些多余数据,只保留基于原时间索引的重采样结果。
原代码
ROD['time'] = pd.to_datetime(ROD['timestamp']) ROD.set_index('time', inplace = True, drop = True) resampled = ROD.resample('60Min',origin='start').agg({'open':'first', 'high':'max', 'low': 'min', 'close': 'last', 'volume':'sum'})
重采样输出示例
open high low close volume time 2020-09-28 10:31:00 333.984985 334.470001 332.779999 333.750000 6482408 2020-09-28 11:31:00 333.760010 334.260010 333.109985 333.469910 4456465 ... 2022-09-23 15:31:00 366.220001 368.040008 366.000000 367.440002 14253081
解决方案
方法1:重采样后过滤时间范围
直接基于原始数据的时间区间,过滤重采样结果:
# 只保留时间在10:30到16:00之间的行(15:59属于15-16点区间) filtered_resampled = resampled[(resampled.index.time >= pd.Timestamp('10:30').time()) & (resampled.index.time < pd.Timestamp('16:00').time())]
或者用原始数据的首尾时间精准过滤:
start = ROD.index.min().floor('H') + pd.Timedelta(minutes=30) end = ROD.index.max().ceil('H') filtered_resampled = resampled[(resampled.index >= start) & (resampled.index < end)]
方法2:调整重采样参数+过滤
修改resample的closed和label参数对齐区间,再过滤时间范围:
# 设置closed='right'、label='right'让区间匹配原始数据时段 resampled = ROD.resample('60Min', closed='right', label='right').agg({ 'open':'first', 'high':'max', 'low': 'min', 'close': 'last', 'volume':'sum' }) # 直接用between_time过滤目标时段 filtered_resampled = resampled.between_time('10:30', '16:00')
方法3:用groupby替代resample
按小时分组后删除无数据的分组:
resampled = ROD.groupby(pd.Grouper(freq='60Min', origin='start')).agg({ 'open':'first', 'high':'max', 'low': 'min', 'close': 'last', 'volume':'sum' }) # 移除全空的行(即无原始数据覆盖的小时区间) filtered_resampled = resampled.dropna(how='all')
内容的提问来源于stack exchange,提问作者svlad
相关产品推荐
相关产品推荐

