如何通过重复值将pandas小时时间序列插值为10分钟粒度数据?
重采样实现固定值重复填充的方案
操作逻辑很简单:先把时间列转为可识别的时间格式,10分钟粒度重采样后调用前向填充方法即可,会自动用前一个小时的原值填充所有间隔点位的空值。
完整可运行代码如下:
import pandas as pd # 构造示例数据集 data = [ ('2014-02-24 16:00:00', 55), ('2014-02-24 17:00:00', 40), ('2014-02-24 18:00:00', 68) ] df = pd.DataFrame(data, columns=['DateTime', 'Value']) # 转换时间列为pandas可识别的datetime格式 df['DateTime'] = pd.to_datetime(df['DateTime']) # 方式1:设置时间列为索引后重采样 resampled_df = df.set_index('DateTime').resample('10T').ffill().reset_index() # 方式2:无需修改索引,直接通过on参数指定时间列 # resampled_df = df.resample('10T', on='DateTime').ffill().reset_index(drop=True)
代码中
10T是pandas时间频率的标准写法,和10Min效果完全一致;ffill()是前向填充方法,会用距离空值最近的上一个非空值做填充,完全匹配直接重复原值的插值需求。
如果需要覆盖最后一个小时的全量10分钟点位(比如示例中18:00之后要生成18:10~18:50的数值),可以先扩展时间边界再做填充:
# 扩展最后一小时边界实现全量覆盖 df.loc[len(df)] = [df['DateTime'].max() + pd.Timedelta(hours=1), pd.NA] resampled_df = df.resample('10T', on='DateTime').ffill().iloc[:-1].reset_index(drop=True)
内容的提问来源于stack exchange,提问作者Fluxy
相关产品推荐
相关产品推荐

