Python环境下含空值Numpy序列基于邻域的异常值检测方案
车道检测序列数据空值前置异常检测方案
我正在进行道路车道检测相关开发,采集到的序列数据中存在标注为-1的空值以及噪声干扰。针对空值我已经编写了如下bridging函数进行插值填补:
def bridging(ar: np.ndarray): v_start = 0 temp = 0 latch = False for e, v in enumerate(ar): if not latch: if v == -1: # enter -1 values latch = True v_start = e temp = ar[e-1] if latch: if v != -1: latch = False v_length = e - v_start + 2 fill_values = np.linspace(temp, v, v_length) ar[v_start:e] = fill_values[1:-1] return ar
但在执行空值填补之前,我需要先基于数值的邻域特征检测并删除异常值。我曾尝试使用Savitzky Golay Filtering实现该需求,但其不支持传入含空值的数据,因此寻求可行的解决方案。
参考示例数据如下,其中加粗标注的数值为需要检出的异常值:
- 序列1:[184, -1, -1, -1, 756, 430, 473, 567, 618, 589, 585, 464, 467, 184, -1, 642, -1, 389, 387, -1, 589, 602, 728, 597, 568, 620, 610, 548, 424, -1, -1, 301, -1, -1, -1, -1, 637]
- 序列2:[ -1, -1, 3686, -1, 3740, 3653, -1, 3656, 3633, -1, -1, 3421, 3389, 3560, -1, -1, -1, 3340, 3313, -1, 3418, 3566, 3643, 3751, 3580, 3686, 3683, 3625, 3515, 3467, -1, 3431, -1, -1, -1, -1, -1]
解决方案
实现思路
无需提前补全空值,采用支持缺失值的鲁棒异常检测逻辑,将检出的异常值统一转为-1后,直接复用你已有的bridging函数做插值即可,完全适配现有处理流程。
核心代码实现
import numpy as np def detect_outliers(ar: np.ndarray, window_size=5, threshold=2.5): ar = ar.copy() # 空值转NaN方便后续邻域统计 ar_nan = np.where(ar == -1, np.nan, ar) seq_len = len(ar_nan) for idx in range(seq_len): current_val = ar_nan[idx] if np.isnan(current_val): continue # 取当前点前后window_size的邻域范围 left_bound = max(0, idx - window_size) right_bound = min(seq_len, idx + window_size + 1) window_vals = ar_nan[left_bound:right_bound] # 邻域全空跳过判断 if np.all(np.isnan(window_vals)): continue # 用中位数绝对偏差(MAD)做鲁棒异常判断,不受极端值和空值影响 window_med = np.nanmedian(window_vals) mad = np.nanmedian(np.abs(window_vals - window_med)) if mad == 0: continue # MAD转换为等效标准差的系数为1.4826 z_score = np.abs(current_val - window_med) / (mad * 1.4826) if z_score > threshold: # 异常值转为-1,后续统一由bridging函数插值 ar[idx] = -1 return ar
使用流程
# 第一步:异常检测,异常值转为-1 processed_seq = detect_outliers(original_seq, window_size=5, threshold=2.5) # 第二步:调用现有插值函数补全所有-1 filled_seq = bridging(processed_seq)
参数调整说明
window_size:根据序列采样频率调整,车道线坐标序列推荐取3~7threshold:阈值越高误检率越低、漏检率越高,推荐在2~3区间内根据实际数据效果调整
内容的提问来源于stack exchange,提问作者TaekYoung
相关产品推荐
相关产品推荐

