You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyOD处理超3500行样本报错求助及大样本分析方案需求

问题排查与解决方案

问题描述

拥有三个样本数据集,约2300行的样本可正常运行PyOD异常检测脚本,但导入超过3500行的样本时脚本报错,报错类型为IndexError(列表索引超出范围),不确定是脚本问题还是PyOD对数据量有上限,需要排查方法及大样本分析方案。

运行脚本

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import matplotlib as matplotlib
from scipy import stats
from pyod.models.abod import ABOD
from pyod.models.cblof import CBLOF
from pyod.models.feature_bagging import FeatureBagging
from pyod.models.hbos import HBOS
from pyod.models.iforest import IForest
from pyod.models.knn import KNN
from pyod.models.lof import LOF

df = pd.read_csv("D:/1/PyOD sample 2.csv")
df.plot.scatter('Location','Value')

from sklearn.preprocessing import MinMaxScaler
scaler = MinMaxScaler(feature_range=(0, 1))
df[['Location','Value']] = scaler.fit_transform(df[['Location','Value']])
df[['Location','Value']].head()

X1 = df['Location'].values.reshape(-1,1)
X2 = df['Value'].values.reshape(-1,1)
X = np.concatenate((X1,X2),axis=1)

random_state = np.random.RandomState(42)
outliers_fraction = 0.05

classifiers = {
    'Angle-based Outlier Detector (ABOD)': ABOD(contamination=outliers_fraction),
    'Cluster-based Local Outlier Factor (CBLOF)':CBLOF(contamination=outliers_fraction,check_estimator=False, random_state=random_state),
    'Feature Bagging':FeatureBagging(LOF(n_neighbors=35),contamination=outliers_fraction,check_estimator=False,random_state=random_state),
    'Histogram-base Outlier Detection (HBOS)': HBOS(contamination=outliers_fraction),
    'Isolation Forest': IForest(contamination=outliers_fraction,random_state=random_state),
    'K Nearest Neighbors (KNN)': KNN(contamination=outliers_fraction),
    'Average KNN': KNN(method='mean',contamination=outliers_fraction)
}

xx , yy = np.meshgrid(np.linspace(0,1 , 200), np.linspace(0, 1, 200))

for i, (clf_name, clf) in enumerate(classifiers.items()):
    clf.fit(X)
    scores_pred = clf.decision_function(X) * -1
    y_pred = clf.predict(X)
    n_inliers = len(y_pred) - np.count_nonzero(y_pred)
    n_outliers = np.count_nonzero(y_pred == 1)
    plt.figure(figsize=(10, 10))
    
    # 修复原脚本的引用问题,避免修改原DataFrame
    dfx = df.copy()
    dfx['outlier'] = y_pred.tolist()
    
    IX1 = np.array(dfx['Location'][dfx['outlier'] == 0]).reshape(-1,1)
    IX2 = np.array(dfx['Value'][dfx['outlier'] == 0]).reshape(-1,1)
    OX1 = dfx['Location'][dfx['outlier'] == 1].values.reshape(-1,1)
    OX2 = dfx['Value'][dfx['outlier'] == 1].values.reshape(-1,1)
    
    print('OUTLIERS : ',n_outliers,'INLIERS : ',n_inliers, clf_name)
    
    # 替换废弃的scipy函数,改用numpy的percentile
    threshold = np.percentile(scores_pred, 100 * outliers_fraction)
    
    Z = clf.decision_function(np.c_[xx.ravel(), yy.ravel()]) * -1
    Z = Z.reshape(xx.shape)
    
    # fill blue map colormap from minimum anomaly score to threshold value
    plt.contourf(xx, yy, Z, levels=np.linspace(Z.min(), threshold, 7),cmap=plt.cm.Blues_r)
        
    # draw red contour line where anomaly score is equal to thresold
    a = plt.contour(xx, yy, Z, levels=[threshold],linewidths=2, colors='red')
        
    # fill orange contour lines where range of anomaly score is from threshold to maximum anomaly score
    plt.contourf(xx, yy, Z, levels=[threshold, Z.max()],colors='orange')
        
    b = plt.scatter(IX1,IX2, c='white',s=20, edgecolor='k')
    c = plt.scatter(OX1,OX2, c='black',s=20, edgecolor='k')
       
    plt.axis('tight')  

    # 修复legend的索引越界问题:检查contour是否生成了有效集合
    legend_handles = []
    legend_labels = []
    if len(a.collections) > 0:
        legend_handles.append(a.collections[0])
        legend_labels.append('learned decision function')
    legend_handles.extend([b, c])
    legend_labels.extend(['inliers','outliers'])
    
    plt.legend(
        legend_handles,
        legend_labels,
        prop=matplotlib.font_manager.FontProperties(size=20),
        loc=2)
    
    plt.xlim((0, 1))
    plt.ylim((0, 1))
    plt.title(clf_name)
    plt.show()

报错堆栈信息

IndexError                                Traceback (most recent call last)
d:\1\PyOD.py in line 69
     66 plt.axis('tight')  
     68 # loc=2 is used for the top left corner 
---> 69 plt.legend(
     70     [a.collections[0], b,c],
     71     ['learned decision function', 'inliers','outliers'],
     72     prop=matplotlib.font_manager.FontProperties(size=20),
     73     loc=2)
     74 plt.xlim((0, 1))
     75 plt.ylim((0, 1))

File c:\Users\z004eeud\AppData\Local\Programs\Python\Python311\Lib\site-packages\matplotlib\pyplot.py:2710, in legend(*args, **kwargs)
   2708 @_copy_docstring_and_deprecators(Axes.legend)
   2709 def legend(*args, **kwargs):
-> 2710     return gca().legend(*args, **kwargs)

File c:\Users\z004eeud\AppData\Local\Programs\Python\Python311\Lib\site-packages\matplotlib\axes\_axes.py:318, in Axes.legend(self, *args, **kwargs)
    316 if len(extra_args):
    317     raise TypeError('legend only accepts two non-keyword arguments')
--> 318 self.legend_ = mlegend.Legend(self, handles, labels, **kwargs)
    319 self.legend_._remove_method = self._remove_legend
    320 return self.legend_

File c:\Users\z004eeud\AppData\Local\Programs\Python\Python311\Lib\site-packages\matplotlib\_api\deprecation.py:454, in make_keyword_only..wrapper(*args, **kwargs)
    448 if len(args) > name_idx:
    449     warn_deprecated(
    450         since, message="Passing the %(name)s %(obj_type)s "
    451         "positionally is deprecated since Matplotlib %(since)s; the "
    452         "parameter will become keyword-only %(removal)s.",
    453         name=name, obj_type=f"parameter of {func.__name__}()")
--> 454 return func(*args, **kwargs)

File c:\Users\z004eeud\AppData\Local\Programs\Python\Python311\Lib\site-packages\matplotlib\legend.py:583, in Legend.__init__(self, parent, handles, labels, loc, numpoints, markerscale, markerfirst, reverse, scatterpoints, scatteryoffsets, prop, fontsize, labelcolor, borderpad, labelspacing, handlelength, handleheight, handletextpad, borderaxespad, columnspacing, ncols, mode, fancybox, shadow, title, title_fontsize, framealpha, edgecolor, facecolor, bbox_to_anchor, bbox_transform, frameon, handler_map, title_fontproperties, alignment, ncol, draggable)
    580 self._alignment = alignment
    582 # init with null renderer
--> 583 self._init_legend_box(handles, labels, markerfirst)
    585 tmp = self._loc_used_default
    586 self._set_loc(loc)

File c:\Users\z004eeud\AppData\Local\Programs\Python\Python311\Lib\site-packages\matplotlib\legend.py:867, in Legend._init_legend_box(self, handles, labels, markerfirst)
    864         text_list.append(textbox._text)
    865         # Create the artist for the legend which represents the
    866         # original artist/handle.
--> 867         handle_list.append(handler.legend_artist(self, orig_handle,
    868                                                  fontsize, handlebox))
    869         handles_and_labels.append((handlebox, textbox))
    871 columnbox = []

File c:\Users\z004eeud\AppData\Local\Programs\Python\Python311\Lib\site-packages\matplotlib\legend_handler.py:130, in HandlerBase.legend_artist(self, legend, orig_handle, fontsize, handlebox)
    106 """
    107 Return the artist that this HandlerBase generates for the given
    108 original artist/handle.
   (...)
    123 
    124 """
    125 xdescent, ydescent, width, height = self.adjust_drawing_area(
    126          legend, orig_handle,
    127          handlebox.xdescent, handlebox.ydescent,
    128          handlebox.width, handlebox.height,
    129          fontsize)
--> 130 artists = self.create_artists(legend, orig_handle,
    131                               xdescent, ydescent, width, height,
    132                               fontsize, handlebox.get_transform())
    134 # create_artists will return a list of artists.
    135 for a in artists:

File c:\Users\z004eeud\AppData\Local\Programs\Python\Python311\Lib\site-packages\matplotlib\legend_handler.py:496, in HandlerRegularPolyCollection.create_artists(self, legend, orig_handle, xdescent, ydescent, width, height, fontsize, trans)
    490 ydata = self.get_ydata(legend, xdescent, ydescent,
    491                        width, height, fontsize)
    493 sizes = self.get_sizes(legend, orig_handle, xdescent, ydescent,
    494                        width, height, fontsize)
--> 496 p = self.create_collection(
    497     orig_handle, sizes,
    498     offsets=list(zip(xdata_marker, ydata)), offset_transform=trans)
    500 self.update_prop(p, orig_handle, legend)
    501 p.set_offset_transform(trans)

File c:\Users\z004eeud\AppData\Local\Programs\Python\Python311\Lib\site-packages\matplotlib\_api\deprecation.py:297, in rename_parameter..wrapper(*args, **kwargs)
    292     warn_deprecated(
    293         since, message=f"The {old!r} parameter of {func.__name__}() "
    294         f"has been renamed {new!r} since Matplotlib {since}; support "
    295         f"for the old name will be dropped %(removal)s.")
    296     kwargs[new] = kwargs.pop(old)
--> 297 return func(*args, **kwargs)

File c:\Users\z004eeud\AppData\Local\Programs\Python\Python311\Lib\site-packages\matplotlib\legend_handler.py:511, in HandlerPathCollection.create_collection(self, orig_handle, sizes, offsets, offset_transform)
    508 @_api.rename_parameter("3.6", "transOffset", "offset_transform")
    509 def create_collection(self, orig_handle, sizes, offsets, offset_transform):
    510     return type(orig_handle)(
--> 511         [orig_handle.get_paths()[0]], sizes=sizes,
    512         offsets=offsets, offset_transform=offset_transform,
    513     )

IndexError: list index out of range

问题根源分析

  1. 报错直接原因:plt.contour生成的轮廓线集合a.collections为空,导致a.collections[0]索引越界。这种情况发生在所有网格点的异常分数都等于阈值时,不会生成任何轮廓线。大样本下数据分布更密集,更容易出现这种极端情况。
  2. 间接诱因:原脚本使用了已废弃的stats.scoreatpercentile函数,在大样本下可能计算出的阈值存在偏差,加剧了轮廓线无法生成的概率。
  3. 脚本隐患:dfx = df是直接引用原DataFrame,会导致原数据被意外修改;绘图时的meshgrid点数过多(200x200),大样本下会增加内存占用和绘图压力。

排查与解决方案

1. 即时修复报错

修改脚本中的legend生成逻辑,先检查轮廓线集合是否为空,再决定是否加入legend:

# 替换原legend代码
legend_handles = []
legend_labels = []
if len(a.collections) > 0:
    legend_handles.append(a.collections[0])
    legend_labels.append('learned decision function')
legend_handles.extend([b, c])
legend_labels.extend(['inliers','outliers'])

plt.legend(
    legend_handles,
    legend_labels,
    prop=matplotlib.font_manager.FontProperties(size=20),
    loc=2)

2. 替换废弃函数

将stats.scoreatpercentile替换为np.percentile,避免大样本下的计算偏差:

# 替换原阈值计算代码
threshold = np.percentile(scores_pred, 100 * outliers_fraction)

3. 脚本优化(大样本适配)

  • 减少meshgrid点数:将np.meshgrid(np.linspace(0,1 , 200), np.linspace(0, 1, 200))中的200改为100或50,降低绘图内存占用。
  • 避免修改原数据:将dfx = df改为dfx = df.copy(),防止原DataFrame被意外修改。
  • 批量绘图或跳过绘图:如果只需要异常检测结果,不需要可视化,可以注释掉绘图相关代码,大幅提升大样本处理速度。

4. PyOD大样本适配建议

PyOD本身对数据量没有严格上限,但部分模型(如ABOD、LOF)在超大数据集下速度较慢,可选择以下优化方案:

  • 使用增量学习模型:如pyod.models.incremental中的增量异常检测器,适合流式大样本。
  • 选择高效模型:优先使用Isolation Forest、HBOS、KNN(带kd-tree优化)等时间复杂度较低的模型。
  • 数据采样:对超大数据集先进行随机采样,再训练模型,后续用全量数据检测异常。

内容的提问来源于stack exchange,提问作者Counterpoint

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 19:07:31