Pandas时间序列重采样遇时间大间隙报错的解决方案咨询
时间序列重采样问题解决方案
原始数据
import pandas as pd import numpy as np ts = pd.Series( [6.22, 6.23, 6.23, 6.24, 6.24, 6.25, np.nan, np.nan, np.nan, np.nan], index=pd.DatetimeIndex( ['2023-08-01 10:31:40.110000', '2023-08-01 10:31:43.110000', '2023-08-01 10:31:46.111000', '2023-08-01 10:31:49.111000', '2023-08-01 10:31:52.111000', '2023-08-01 10:31:55.117000', '2023-08-01 10:31:58.112000', '2023-08-01 10:32:01.112000', '2023-08-01 10:32:04.117000', '2023-08-01 10:34:07.095000'], dtype='datetime64[ns]', name='exchange_time', freq=None ) )
问题描述
- 使用
ts.resample('6S', closed='left', label='right').apply(lambda x: x.iloc[-1])时,因时间索引存在较大间隙,空采样区间调用x.iloc[-1]会直接报错 last()方法无法满足需求:要求2023-08-01 10:32:00对应的重采样结果保留原始数据的NaN值,但last()会忽略空值或无法正确匹配区间逻辑
解决方案
通过在自定义函数中先判断采样区间是否为空,既解决报错问题又严格保留原始空值:
result = ts.resample('6S', closed='left', label='right').apply(lambda x: x.iloc[-1] if not x.empty else np.nan)
代码说明
not x.empty判断当前采样区间是否包含数据,空区间直接返回NaN,彻底解决时间间隙导致的报错- 非空区间取最后一个元素,无论该元素是数值还是NaN,严格保留原始数据的空值状态
- 采样参数
closed='left', label='right'保持原逻辑:区间左闭右开,用区间右端点作为结果标签
验证结果中,2023-08-01 10:32:00对应的区间包含2023-08-01 10:31:58.112000的NaN值,因此结果为NaN,完全符合需求。
内容的提问来源于stack exchange,提问作者tesla1060
相关产品推荐
相关产品推荐

