基于直方图孤立边缘检测时间序列异常值的代码误判问题排查
修正野外传感器时间序列异常值检测逻辑
问题根源
之前的代码只靠「计数<50」标记孤立边缘,漏掉了两个核心判断条件:邻接区间计数为0、区间位于直方图首尾(极端位置),才会把靠近主分布的43、46这类低计数区间误判成异常。
修正后的判断规则
只有同时满足以下三个条件,才标记为通信故障导致的极端异常边缘:
- 区间计数占总样本量**<1%**(替代固定的50,适配不同数据规模)
- 左右邻接区间至少有一个计数为0(首尾区间只需单侧邻接为0)
- 该区间是直方图的最左端或最右端(对应极端异常的位置)
修正代码实现
import numpy as np import matplotlib.pyplot as plt def detect_isolated_outlier_bins(data, bin_count=20): counts, bins = np.histogram(data, bins=bin_count) isolated_bins = [] total_samples = len(data) # 按占比设阈值,确保异常占比<1% count_threshold = total_samples * 0.01 for i in range(len(counts)): current_count = counts[i] if current_count < count_threshold: # 判断邻接区间是否为空 left_empty = (i == 0) or (counts[i-1] == 0) right_empty = (i == len(counts)-1) or (counts[i+1] == 0) # 必须是首尾区间,且至少一侧邻接为空 if (left_empty or right_empty) and (i == 0 or i == len(counts)-1): isolated_bins.append(bins[i]) return isolated_bins # 模拟你的数据场景:0是真实异常,43/46是靠近主分布的低计数 data = np.concatenate([ np.zeros(30), # 真实异常,占比<1% np.random.randint(40, 50, 1000), # 主分布数据 np.array([43]*40), # 靠近主分布的低计数 np.array([46]*45) # 靠近主分布的低计数 ]) # 检测异常边缘 isolated_bins = detect_isolated_outlier_bins(data) print("检测到的孤立异常边缘:", isolated_bins) # 剔除异常样本 filtered_data = data[~np.isin(data, isolated_bins)] # 可视化验证效果 plt.hist(data, bins=20, alpha=0.5, label='原始数据') plt.hist(filtered_data, bins=20, alpha=0.5, label='过滤后数据') plt.legend() plt.show()
关键优化点
- 用占比阈值替代固定数值50,不管数据量大小都能保证异常占比<1%
- 加入邻接区间空值判断,确保标记的是真正“孤立”的边缘
- 限定仅首尾区间会被标记,精准定位通信故障导致的极端异常
内容的提问来源于stack exchange,提问作者Mainland
相关产品推荐
相关产品推荐

