You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按最小时间间隔高效筛选Pandas时间序列/数据帧?

高效筛选Pandas中满足最小时间间隔的时间戳数据

我有一个带不规则间隔时间戳的Pandas DataFrame/Series,需要筛选出**行间最小时间间隔不小于指定值(比如20ms)**的数据,间隔允许更大。以下是筛选前后的示例:

# 筛选前
317   2022-12-31 00:00:00.360
318   2022-12-31 00:00:00.364
319   2022-12-31 00:00:00.368
320   2022-12-31 00:00:00.372
321   2022-12-31 00:00:00.376
322   2022-12-31 00:00:00.380
323   2022-12-31 00:00:00.384
324   2022-12-31 00:00:00.388
325   2022-12-31 00:00:00.392
326   2022-12-31 00:00:00.396
327   2022-12-31 00:00:00.414
328   2022-12-31 00:00:00.416
329   2022-12-31 00:00:00.420
330   2022-12-31 00:00:00.425
331   2022-12-31 00:00:00.428
332   2022-12-31 00:00:00.432
333   2022-12-31 00:00:00.438

# 筛选后(最小间隔20ms)
317   2022-12-31 00:00:00.360
320   2022-12-31 00:00:00.372
325   2022-12-31 00:00:00.392
327   2022-12-31 00:00:00.414
333   2022-12-31 00:00:00.438

当前我用简单for循环实现需求,但数据量极大时效率极低:

res = [timestamps[0]]
min_delta = pd.Timedelta('20ms')
for dt in timestamps[1:]:
    if dt - res[-1] >= min_delta:
        res.append(dt)

尝试过resample、diff等向量化方法但没得到预期结果,求高效的可行方案。


高效解决方案:向量化累积筛选

由于需求是当前时间戳与最后保留的时间戳比较,而非与前一行比较,直接用diff()无法实现。可以用NumPy底层的循环实现,比Python层级循环快数倍:

import pandas as pd
import numpy as np

def filter_min_interval(timestamps, min_delta):
    # 转换为numpy数组提升运算效率
    ts_np = timestamps.to_numpy()
    min_delta_ns = min_delta.value  # 转为纳秒数值
    
    # 初始化保留标记
    keep = np.zeros(len(ts_np), dtype=bool)
    keep[0] = True
    last_kept_ts = ts_np[0].value
    
    # 利用NumPy的C层级循环遍历
    for i in range(1, len(ts_np)):
        current_ts = ts_np[i].value
        if current_ts - last_kept_ts >= min_delta_ns:
            keep[i] = True
            last_kept_ts = current_ts
    
    return timestamps[keep]

进阶加速:NumJIT编译循环

如果可以安装numba库,用JIT编译循环能达到接近纯C的性能,适合超大规模数据集:

from numba import njit
import pandas as pd
import numpy as np

@njit
def _filter_numba(ts_values, min_delta_ns):
    keep = np.zeros(len(ts_values), dtype=np.bool_)
    keep[0] = True
    last_kept = ts_values[0]
    for i in range(1, len(ts_values)):
        if ts_values[i] - last_kept >= min_delta_ns:
            keep[i] = True
            last_kept = ts_values[i]
    return keep

def filter_min_interval_numba(timestamps, min_delta):
    # 将时间戳转为纳秒整数数组
    ts_values = timestamps.to_numpy().astype(np.int64)
    min_delta_ns = min_delta.value
    keep = _filter_numba(ts_values, min_delta_ns)
    return timestamps[keep]

使用示例

# 假设timestamps是Pandas Series,类型为datetime64
timestamps = pd.Series(pd.to_datetime([
    '2022-12-31 00:00:00.360', '2022-12-31 00:00:00.364',
    '2022-12-31 00:00:00.368', '2022-12-31 00:00:00.372',
    '2022-12-31 00:00:00.376', '2022-12-31 00:00:00.380',
    '2022-12-31 00:00:00.384', '2022-12-31 00:00:00.388',
    '2022-12-31 00:00:00.392', '2022-12-31 00:00:00.396',
    '2022-12-31 00:00:00.414', '2022-12-31 00:00:00.416',
    '2022-12-31 00:00:00.420', '2022-12-31 00:00:00.425',
    '2022-12-31 00:00:00.428', '2022-12-31 00:00:00.432',
    '2022-12-31 00:00:00.438'
]))

min_delta = pd.Timedelta('20ms')
filtered = filter_min_interval(timestamps, min_delta)
print(filtered)

内容的提问来源于stack exchange,提问作者Simon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 15:54:56