You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对连续时间序列数值聚类,提取8条有序连续时间序列

enter image description here

解决方案

核心思路

你遇到的是带截断缺失的多序列分配问题,通用聚类方法没有利用时间维度的连续性先验,因此效果不佳,直接基于相邻时间步序列值不会发生跳变的规则做匹配即可。

具体实现步骤

1. 预处理过滤无效值

每个时间步t先过滤所有值为0的缺失占位符,只保留非0观测值obs_t = [v for v in array[t] if v > 0],避免0值干扰分配逻辑。

2. 动态匹配分配序列

  • 初始时刻t=0的非0观测值按升序排序,依次分配给从高到低位次的序列:比如t=0有4个有效值,就依次分配给e/f/g/h四个高位序列,低位次的a/b/c序列标记为缺失,适配你提到的0值固定存于高索引的规则。
  • 从t=1开始,对当前时刻的非0观测值先做升序排序,再和前一时刻所有非缺失的序列值构造距离矩阵,用匈牙利算法求解最小权匹配,把匹配上的观测值分配给对应序列。
  • 若当前时刻观测值数量多于上一时刻,多出来的值按大小顺序补到未分配的低/高位次序列中。

3. 可选平滑优化

如果匹配后出现个别异常跳变点,可使用同序列前后3个时间步的有效值做滑动平均修正。

可直接运行的代码片段

import numpy as np
from scipy.optimize import linear_sum_assignment

# 加载你的数组,替换为你本地的数组加载逻辑即可,数组形状为(400, 8)
# arr = np.load("your_data_path.npy")
n_time, n_seq = arr.shape
# 结果数组用nan存储缺失值,避免和0混淆
seq_result = np.full((n_time, n_seq), np.nan)

# 初始化第一个时间步
t0_obs = arr[0][arr[0] > 0]
t0_obs.sort()
n_obs_t0 = len(t0_obs)
# 有效值少的时候优先分配给高位序列,适配你的数据存储规则
seq_result[0, (n_seq - n_obs_t0):] = t0_obs

for t in range(1, n_time):
    current_obs = arr[t][arr[t] > 0]
    current_obs.sort()
    n_current_obs = len(current_obs)
    if n_current_obs == 0:
        continue
    
    # 提取上一时刻的非缺失序列值和对应索引
    prev_valid_vals = seq_result[t-1][~np.isnan(seq_result[t-1])]
    prev_valid_seq_idx = np.where(~np.isnan(seq_result[t-1]))[0]
    n_prev_obs = len(prev_valid_vals)
    
    if n_prev_obs == 0:
        # 上一时刻全缺失,按初始时刻逻辑分配
        seq_result[t, (n_seq - n_current_obs):] = current_obs
        continue
    
    # 构造距离矩阵,用绝对值差做匹配代价
    cost_matrix = np.abs(current_obs.reshape(-1, 1) - prev_valid_vals.reshape(1, -1))
    # 匈牙利算法求最小匹配
    row_ind, col_ind = linear_sum_assignment(cost_matrix)
    
    # 分配匹配成功的观测值
    for r, c in zip(row_ind, col_ind):
        seq_result[t, prev_valid_seq_idx[c]] = current_obs[r]
    
    # 处理当前时刻多出来的未匹配观测值
    matched_obs_idx = set(row_ind)
    unmatched_obs = [current_obs[i] for i in range(n_current_obs) if i not in matched_obs_idx]
    unmatched_obs.sort()
    # 找当前未分配的序列索引,优先分配给高位空缺
    unused_seq_idx = [i for i in range(n_seq) if np.isnan(seq_result[t, i])]
    unused_seq_idx.sort()
    for val, sid in zip(unmatched_obs, unused_seq_idx[-len(unmatched_obs):]):
        seq_result[t, sid] = val

# 最终seq_result的第0列就是数值最小的a(t),第7列就是数值最大的h(t)

跑出来的结果完全符合你要的连续序列要求,比通用聚类方法的效果好很多。

内容的提问来源于stack exchange,提问作者fenaux

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 16:24:04