Python中仅对时间间隔≤10分钟的时间序列重采样与插值优化方案
高效处理带间隙时间序列的重采样与插值方案
需求说明
现有10分钟间隔采样的时间序列数据,但存在时间间隙。需仅对相邻时间戳间隔≤10分钟的连续区间,按1秒频率重采样并执行线性插值;相邻间隔超过10分钟的区间直接舍弃(示例中仅保留2023-01-01 00:00:00至2023-01-01 00:20:00、2023-01-01 00:40:00至2023-01-01 00:50:00这两个区间)。当前通过for循环实现,需替换为无需迭代DataFrame的高效方案。
示例数据
import pandas as pd data = { 'timestamp': pd.to_datetime(['2023-01-01 00:00:00', '2023-01-01 00:10:00', '2023-01-01 00:20:00', '2023-01-01 00:40:00', '2023-01-01 00:50:00', '2023-01-01 01:10:00']), 'value': [10, 20, 30, 50, 60, 80] } df = pd.DataFrame(data)
原循环实现代码
df['gap'] = df['timestamp'].diff() > pd.Timedelta(minutes=10) sections = [] current_section = [] for index, row in df.iterrows(): if row['gap']: if current_section: sections.append(current_section) current_section = [] current_section.append(row) sections.append(current_section) resampled_data = [] for section in sections: section_df = pd.DataFrame(section) if len(section_df) > 1: start_time = section_df.iloc[0]['timestamp'] end_time = section_df.iloc[-1]['timestamp'] resampled = section_df.set_index('timestamp').resample('1S').interpolate(method='linear') resampled = resampled.loc[start_time:end_time] resampled_data.append(resampled) resampled_data = pd.concat(resampled_data).drop(columns='gap')
高效无迭代实现方案
利用pandas的分组功能,通过标记连续区间的分组ID,实现批量处理:
import pandas as pd # 1. 标记相邻时间戳是否存在超过10分钟的间隙 df['is_gap'] = df['timestamp'].diff() > pd.Timedelta(minutes=10) # 2. 生成连续区间的分组ID:间隙出现时分组ID递增 df['group_id'] = df['is_gap'].cumsum() # 3. 过滤掉仅含单个数据点的分组(无法插值),然后按分组处理 resampled_result = ( df.groupby('group_id') .filter(lambda x: len(x) > 1) # 保留至少2个数据点的分组 .set_index('timestamp') .groupby('group_id') .resample('1S') .interpolate(method='linear') .drop(columns=['is_gap', 'group_id']) # 移除辅助列 .reset_index(level=0, drop=True) # 移除group_id索引 ) # 结果预览 print(resampled_result.head())
方案说明
- 用
diff()和cumsum()生成分组ID,自动将连续无间隙的区间归为同一组,替代手动循环划分区间的逻辑 - 通过
groupby()+resample()的组合,批量完成每个区间的重采样与插值操作 - 全程基于pandas矢量化操作,避免了迭代DataFrame的性能损耗,代码更简洁易维护
内容的提问来源于stack exchange,提问作者crx91
相关产品推荐
相关产品推荐

