使用LIME解释欺诈检测深度神经网络时explain_instance报索引错误
LIME解释欺诈交易分类模型时触发索引越界问题
问题复现
搭建欺诈交易分类深度神经网络后,调用interpretor.explain_instance()开展模型预测结果解释时触发报错,复现代码如下:
import lime from lime import lime_tabular interpretor = lime_tabular.LimeTabularExplainer( training_data=x_train_scaled, feature_names=X_train.columns, mode='classification' ) exp = interpretor.explain_instance( data_row=x_test_scaled[:1], ##new data predict_fn=model.predict,num_features=11 ) xp.show_in_notebook(show_table=True)
运行后抛出核心报错:
IndexError: index 1 is out of bounds for axis 1 with size 1
错误触发位置为lime_base.py中读取neighborhood_labels对应标签列的语句,提示索引1超出轴1的有效范围,该轴维度大小仅为1。
报错原因
- 核心原因是预测函数输出格式不符合LIME要求:二分类场景下如果模型最后一层用sigmoid激活,
model.predict仅返回正类的预测概率,输出数组shape为(样本数, 1),但LIME分类模式要求预测函数必须返回所有类别的概率值,二分类场景需要shape为(样本数, 2)的输出(分别对应负类、正类概率),LIME尝试读取索引为1的正类概率列时,因数组第二维长度只有1触发越界。 - 代码存在两处笔误:一是最后调用结果展示方法时,误将解释结果变量名
exp写为xp;二是传入待解释样本时用切片x_test_scaled[:1]得到的是带batch维度的二维数组,LIME要求单样本输入为一维数组。
修复方法
- 自定义预测包装函数,将模型单通道的sigmoid输出拼接为两类概率输出,适配LIME的格式要求
- 修正代码笔误,传入一维格式的单样本数据
修复后的可运行代码:
import lime from lime import lime_tabular import numpy as np # 包装预测函数,输出格式适配LIME要求 def predict_adapter(data): pos_prob = model.predict(data, verbose=0) # 拼接负类概率,最终输出shape为(输入样本数, 2) return np.concatenate([1 - pos_prob, pos_prob], axis=1) interpretor = lime_tabular.LimeTabularExplainer( training_data=x_train_scaled, feature_names=X_train.columns, mode='classification', class_names=['正常交易', '欺诈交易'] # 可选配置,指定类别名后展示更直观 ) # 传入一维单样本,调用包装后的预测函数 exp = interpretor.explain_instance( data_row=x_test_scaled[0], predict_fn=predict_adapter, num_features=11 ) # 修正变量名,展示解释结果 exp.show_in_notebook(show_table=True)
注:如果你的分类模型最后一层用softmax激活、直接输出所有类别概率(二分类下输出shape为(n,2)),无需包装预测函数,直接传入
model.predict即可,仅需修正两处笔误就能正常运行。
内容的提问来源于stack exchange,提问作者Gourab
相关产品推荐
相关产品推荐

