You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于直方图孤立边缘检测时间序列异常值的代码误判问题排查

修正野外传感器时间序列异常值检测逻辑

问题根源

之前的代码只靠「计数<50」标记孤立边缘,漏掉了两个核心判断条件:邻接区间计数为0、区间位于直方图首尾(极端位置),才会把靠近主分布的43、46这类低计数区间误判成异常。

修正后的判断规则

只有同时满足以下三个条件,才标记为通信故障导致的极端异常边缘:

  • 区间计数占总样本量**<1%**(替代固定的50,适配不同数据规模)
  • 左右邻接区间至少有一个计数为0(首尾区间只需单侧邻接为0)
  • 该区间是直方图的最左端或最右端(对应极端异常的位置)

修正代码实现

import numpy as np
import matplotlib.pyplot as plt

def detect_isolated_outlier_bins(data, bin_count=20):
    counts, bins = np.histogram(data, bins=bin_count)
    isolated_bins = []
    total_samples = len(data)
    # 按占比设阈值,确保异常占比<1%
    count_threshold = total_samples * 0.01

    for i in range(len(counts)):
        current_count = counts[i]
        if current_count < count_threshold:
            # 判断邻接区间是否为空
            left_empty = (i == 0) or (counts[i-1] == 0)
            right_empty = (i == len(counts)-1) or (counts[i+1] == 0)
            # 必须是首尾区间,且至少一侧邻接为空
            if (left_empty or right_empty) and (i == 0 or i == len(counts)-1):
                isolated_bins.append(bins[i])
    
    return isolated_bins

# 模拟你的数据场景:0是真实异常,43/46是靠近主分布的低计数
data = np.concatenate([
    np.zeros(30),  # 真实异常,占比<1%
    np.random.randint(40, 50, 1000),  # 主分布数据
    np.array([43]*40),  # 靠近主分布的低计数
    np.array([46]*45)   # 靠近主分布的低计数
])

# 检测异常边缘
isolated_bins = detect_isolated_outlier_bins(data)
print("检测到的孤立异常边缘:", isolated_bins)

# 剔除异常样本
filtered_data = data[~np.isin(data, isolated_bins)]

# 可视化验证效果
plt.hist(data, bins=20, alpha=0.5, label='原始数据')
plt.hist(filtered_data, bins=20, alpha=0.5, label='过滤后数据')
plt.legend()
plt.show()

关键优化点

  1. 用占比阈值替代固定数值50,不管数据量大小都能保证异常占比<1%
  2. 加入邻接区间空值判断,确保标记的是真正“孤立”的边缘
  3. 限定仅首尾区间会被标记,精准定位通信故障导致的极端异常

内容的提问来源于stack exchange,提问作者Mainland

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 08:35:18