如何在numpy数组指定列中查找数值序列并提取对应特征
纯Numpy实现方案
该方案支持任意长度的查询序列,无额外依赖,兼容性适配老旧版本Numpy:
import numpy as np def find_matching_sequences(label_arr, query_seq): n = len(label_arr) k = len(query_seq) # 边界处理:查询序列长度超过原数组直接返回空 if k > n: return np.array([]), [] # 构造所有长度为k的滑动窗口索引 window_idx = np.arange(k) + np.arange(n - k + 1)[:, None] # 取出所有窗口的label值与查询序列全匹配比对 matches = (label_arr[window_idx] == query_seq).all(axis=1) # 提取匹配的起始索引 start_indices = np.where(matches)[0] # 生成完整匹配的索引列表 full_indices = [list(range(s, s + k)) for s in start_indices] return start_indices, full_indices # 你的场景调用示例 v = np.array([[121,1], [131,1], [113,1], [131,1], [223,1], [242,1], [212,1], [131,1], [113,1], [131,1]]) label_col = v[:, 0] sequence = [131,113,131] start_indices, full_indices = find_matching_sequences(label_col, sequence) # 输出结果:[1 7],[[1, 2, 3], [7, 8, 9]] print("匹配起始索引:", start_indices) print("完整匹配索引列表:", full_indices) # 提取对应feature的示例 feature_arr = ['a','b','c','d','e','f','g','h','i','j'] matched_features = [[feature_arr[i] for i in idx_list] for idx_list in full_indices] # 输出结果:[['b','c','d'], ['h','i','j']] print("匹配的feature序列:", matched_features)
若你使用的Numpy版本过老不支持上述广播写法,可替换为
np.lib.stride_tricks.as_strided构造滑动窗口,逻辑完全一致。
Pandas实现思路(可选)
如果允许使用pandas,可通过滚动窗口实现,逻辑更直观:
import pandas as pd df = pd.DataFrame(v, columns=['label', 'other']) df['feature'] = ['a','b','c','d','e','f','g','h','i','j'] k = len(sequence) # 滚动窗口全匹配判断 match_mask = df['label'].rolling(k).apply(lambda x: (x == sequence).all()).fillna(0).astype(bool) # 匹配起始索引 = 匹配位置往前推k-1位 start_indices = df.index[match_mask] - k + 1
内容的提问来源于stack exchange,提问作者Kunal Shah
相关产品推荐
相关产品推荐

