You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas对不同时间戳索引的序列重采样与合并?

基于不规则时间索引的Pandas重采样方案

需求一:高分辨率数据下采样(按低分辨率区间取均值)

无需循环,可通过pd.cut将高分辨率时间索引映射到低分辨率的区间,再用groupby计算区间均值,完全依托Pandas内置函数实现高效处理。

实现代码

import pandas as pd
import numpy as np

# 高分辨率数据
index1 = pd.date_range('1/1/2000', periods=30, freq='S')
d = np.array([1 for i in range(30)])
d[12:23] = 8
hi_res_data = pd.Series(d, index=index1)

# 低分辨率数据
index2 = pd.date_range('1/1/2000 00:00:00.5', periods=6, freq='5S')
low_res_data = pd.Series(range(100,106), index=index2)

# 生成区间边界:扩展右端点避免遗漏最后一段数据
bins = list(index2) + [index2[-1] + pd.Timedelta(seconds=5)]
# 将高分辨率索引分配到对应区间
hi_res_interval = pd.cut(hi_res_data.index, bins=bins, include_lowest=True)

# 按区间分组计算均值,结果对齐低分辨率时间索引
downsampled_data = hi_res_data.groupby(hi_res_interval).mean()
downsampled_data.index = index2

# 输出与手动实现的desired_data(忽略第一个0)一致的结果
print(downsampled_data.values)

需求二:低分辨率数据上采样到高分辨率索引并插值

利用reindex将低分辨率数据对齐到高分辨率时间索引,再通过interpolate的时间插值能力适配不规则间隔,保留原始数据的时间特性。

实现代码

# 上采样并使用时间插值法适配非规则时间序列
upsampled_data = low_res_data.reindex(hi_res_data.index).interpolate(method='time')

# 查看上采样插值结果
print(upsampled_data.head(10))

关键说明

  • method='time'会根据时间间隔的长短自动计算线性插值,完美适配非规则时间索引的需求;
  • 若需要其他插值逻辑,可替换method参数为'linear'(线性插值)、'nearest'(最近邻插值)等。

内容的提问来源于stack exchange,提问作者Dakiltedyaksman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 05:13:27