如何在Pandas DataFrame中对非均匀时间列进行双倍采样并插值数据?
在非均匀间隔的时间序列中插入中间点并插值
问题描述
现有如下结构的Pandas DataFrame:
>>> pd.DataFrame({'time':[1,2,3], 'data':[10,20,30]}) time data 0 1 10 1 2 20 2 3 30
需要在每一对相邻时间点之间插入中间行(时间为两点的平均值),并对data列做线性插值。预期结果如下:
time data 0 1.0 10 1 1.5 15 2 2.0 20 3 2.5 25 4 3.0 30
注意:time列不一定是均匀间隔的,例如以下数据集同样需要支持:
>>> pd.DataFrame({'time':[1,2.1,3], 'data':[10,20,30]}) time data 0 1.0 10 1 2.1 20 2 3.0 30
由于时间间隔不均匀,resample方法不适用。
解决方案
可以通过生成包含原时间点和中间点的新时间序列,再结合线性插值实现需求,以下是两种简洁的实现方式:
方法一:合并时间点后插值
import pandas as pd import numpy as np # 加载你的数据 df = pd.DataFrame({'time':[1,2.1,3], 'data':[10,20,30]}) # 计算所有相邻时间点的中间值,与原时间数组合并后排序 mid_times = (df['time'].iloc[:-1] + df['time'].iloc[1:]) / 2 new_time = np.sort(np.concatenate([df['time'], mid_times])) # 创建新DataFrame并对缺失的data列做线性插值 new_df = pd.DataFrame({'time': new_time}).merge(df, on='time', how='left').interpolate(method='linear') print(new_df)
方法二:重新索引后插值
import pandas as pd import numpy as np df = pd.DataFrame({'time':[1,2.1,3], 'data':[10,20,30]}) # 生成包含原时间点和中间点的排序后的时间序列 new_time = np.concatenate([df['time'], (df['time'].shift() + df['time'])/2]).dropna().sort_values() # 将原数据按time设为索引,重新索引到新时间序列后插值,最后重置索引 new_df = df.set_index('time').reindex(new_time).interpolate(method='linear').reset_index().rename(columns={'index':'time'}) print(new_df)
说明
两种方法的核心逻辑一致:
- 生成包含原始时间点和所有相邻时间中间点的新时间序列
- 基于新时间序列重构DataFrame,对缺失的
data值执行线性插值
这种方式不受时间间隔是否均匀的限制,完全适配你的需求。
内容的提问来源于stack exchange,提问作者Morten Nissov
相关产品推荐
相关产品推荐

