You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

高效识别多时间序列中符合阈值条件的序列

高效筛选满足阈值条件的时间序列方案

问题描述

我拥有10000个时间序列(每个包含3至10000个数据点),每个序列对应一个唯一阈值。需要找出其中存在至少一个数据点满足小于/大于/等于阈值条件的序列。

示例

threshold_data = [
    # Name Threshold data-points..
    ['ds1', 90,    91, 92, 95],
    ['ds2', 85,    91, 84, 95],
]
  • 操作符为<时,预期输出ds2(其包含小于阈值85的84)
  • 操作符为>时,两个序列均需返回
  • 操作符为==时,无结果

现有方案的痛点

之前尝试过以下方式,但均存在明显缺陷:

  • Pandas逐列循环比较:数据点达10000时效率极低
  • 行内比较数据点子集与阈值:Pandas无法匹配维度
  • 复制阈值生成同尺寸DataFrame:内存占用过大
  • 按阈值分组后用Python循环:分组过多导致速度缓慢

高效实现方案

方案1:Pandas重塑表结构+批量比较

通过melt将宽表转为窄表,利用矢量化操作批量比较,避免循环。

import pandas as pd

threshold_data = [
    ['ds1', 90, 91, 92, 95],
    ['ds2', 85, 91, 84, 95],
]

COL_NAME, COL_THRESHOLD = 'Name', 'Threshold'
# 自动适配动态列数构建DataFrame
df = pd.DataFrame(threshold_data)
df.columns = [COL_NAME, COL_THRESHOLD] + [f't{i}' for i in range(1, len(df.columns)-1)]

# 重塑数据:将所有时间点列转为单行记录
df_melted = df.melt(id_vars=[COL_NAME, COL_THRESHOLD], var_name='time_point', value_name='data')

def filter_sequences(op):
    if op == '<':
        mask = df_melted['data'] < df_melted[COL_THRESHOLD]
    elif op == '>':
        mask = df_melted['data'] > df_melted[COL_THRESHOLD]
    elif op == '==':
        mask = df_melted['data'] == df_melted[COL_THRESHOLD]
    else:
        raise ValueError("仅支持<、>、==三种操作符")
    
    # 分组判断每组是否存在符合条件的记录
    return df_melted[mask].groupby(COL_NAME).size().index.tolist()

# 测试用例
print(filter_sequences('<'))  # 输出: ['ds2']
print(filter_sequences('>'))  # 输出: ['ds1', 'ds2']
print(filter_sequences('==')) # 输出: []

优势:利用Pandas矢量化操作,避免循环开销;内存占用远低于复制阈值生成大表。

方案2:Pandas apply结合numpy矢量化判断

无需修改表结构,直接对每行调用numpy的any方法判断是否存在满足条件的数据点。

import pandas as pd
import numpy as np

threshold_data = [
    ['ds1', 90, 91, 92, 95],
    ['ds2', 85, 91, 84, 95],
]

COL_NAME, COL_THRESHOLD = 'Name', 'Threshold'
df = pd.DataFrame(threshold_data)
df.columns = [COL_NAME, COL_THRESHOLD] + [f't{i}' for i in range(1, len(df.columns)-1)]

# 提取所有数据点列
data_cols = df.columns.drop([COL_NAME, COL_THRESHOLD])

def filter_with_apply(op):
    def check_row(row):
        threshold = row[COL_THRESHOLD]
        data = row[data_cols].values
        if op == '<':
            return np.any(data < threshold)
        elif op == '>':
            return np.any(data > threshold)
        elif op == '==':
            return np.any(data == threshold)
        return False
    
    # 筛选符合条件的序列名称
    return df[df.apply(check_row, axis=1)][COL_NAME].tolist()

# 测试用例
print(filter_with_apply('<'))  # ['ds2']
print(filter_with_apply('>'))  # ['ds1', 'ds2']
print(filter_with_apply('==')) # []

优势:保留原始表结构,numpy的any操作底层为C实现,比纯Python循环快数倍。

方案3:纯numpy处理(超大数据量场景)

直接用numpy数组操作,底层效率最高,适合数据量极大的场景。

import numpy as np

threshold_data = [
    ['ds1', 90, 91, 92, 95],
    ['ds2', 85, 91, 84, 95],
]

# 分离名称、阈值和数据点
names = np.array([row[0] for row in threshold_data])
thresholds = np.array([row[1] for row in threshold_data])
# 补全不同长度的序列(NaN不影响any判断)
max_len = max(len(row)-2 for row in threshold_data)
data = np.array([row[2:] + [np.nan]*(max_len - len(row)+2) for row in threshold_data])

def filter_numpy(op):
    if op == '<':
        mask = np.any(data < thresholds[:, np.newaxis], axis=1)
    elif op == '>':
        mask = np.any(data > thresholds[:, np.newaxis], axis=1)
    elif op == '==':
        mask = np.any(data == thresholds[:, np.newaxis], axis=1)
    else:
        raise ValueError("仅支持<、>、==三种操作符")
    return names[mask].tolist()

# 测试用例
print(filter_numpy('<'))  # ['ds2']
print(filter_numpy('>'))  # ['ds1', 'ds2']
print(filter_numpy('==')) # []

优势:numpy底层为C优化,速度最快;内存控制比Pandas更灵活。

内容的提问来源于stack exchange,提问作者Aaron Digulla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 17:35:16