SHAP Deep Explainer解释LSTM模型时报无masker属性错误如何解决
问题描述
我们使用Keras构建LSTM模型,实现代码如下:
import keras from keras.models import Sequential from keras.layers import Dense from keras.layers import LSTM #make LSTM model architecture model2 = Sequential() model2.add(LSTM(100, return_sequences = True)) model2.add(LSTM(50, return_sequences = True)) model2.add(LSTM(10)) model2.add(Dense(1)) model2.compile(loss='mae', optimizer='adam')
上述模型已完成训练且可正常运行,现需要使用SHAP对该LSTM模型的输出进行可解释性分析。
我们按照如下方式调用SHAP:
import shap explainer = shap.DeepExplainer(model2,x_train_appended) shap_values = explainer(x_train_appended)
执行上述代码时抛出如下错误:
WARNING:tensorflow:Layers in a Sequential model should only have a single input tensor, but we receive a <class 'list'> input: [<tf.Tensor: shape=(49586, 1, 23), dtype=float32, numpy= array([[[0.40824828, 0.02564103, 0.03370786, ..., 0.4494382 , 0.43333334, 0.59210527]], [[0. , 0.06410257, 0.05617978, ..., 0.4494382 , 0.43333334, 0.59210527]], [[0.5400617 , 0.06410257, 0.06741573, ..., 0.4494382 , 0.43333334, 0.59210527]], ..., [[0.5400617 , 0.01282051, 0.05617978, ..., 0.07865169, 0.01111111, 0.05263158]], [[0. , 0.02564103, 0.05617978, ..., 0.07865169, 0.01111111, 0.05263158]], [[0. , 0.02564103, 0.05617978, ..., 0.07865169, 0.01111111, 0.05263158]]], dtype=float32)>] Consider rewriting this model with the Functional API. Traceback (most recent call last): File "", line 3, in shap_values = explainer(x_train_appended) File "/home/kiton/.local/lib/python3.8/site-packages/shap/explainers/_explainer.py", line 207, in call if issubclass(type(self.masker), maskers.OutputComposite) and len(args)==2: AttributeError: 'Deep' object has no attribute 'masker'
问题原因
这个报错是两个问题叠加导致的:
- SHAP版本兼容问题:新版SHAP(0.40.0及以上版本)重构了Explainer的调用逻辑,旧版
DeepExplainer的直接实例化+调用方式和新版的masker机制不兼容,才会抛出'Deep' object has no attribute 'masker'的错误。 - 输入格式问题:TensorFlow/Keras的Sequential模型要求输入为单个张量,SHAP内部处理时将输入封装成了列表结构,触发了TensorFlow的输入格式警告。
解决方案
按以下步骤调整即可正常运行:
- 不要传入全量训练集作为背景样本,从训练集中随机抽取100-200个样本作为背景集即可,既减少计算量,也能避免全量数据带来的格式处理异常。
- 调用
shap_values()方法计算SHAP值,不要直接把explainer实例当函数调用,适配DeepExplainer的原生接口。 - 确保输入数据是numpy数组格式,不要传入tf.Tensor类型的输入。
调整后的代码如下:
import shap import numpy as np # 从训练集随机抽取100个样本作为背景数据 background = x_train_appended[np.random.choice(x_train_appended.shape[0], 100, replace=False)] # 初始化DeepExplainer explainer = shap.DeepExplainer(model2, background) # 调用shap_values方法计算目标数据的SHAP值,按需传入要解释的样本量,全量计算速度会很慢 shap_values = explainer.shap_values(x_train_appended[:1000])
如果还是出现Sequential模型输入的警告,可以把模型用Functional API重写一遍,结构和原模型完全一致即可,不会影响已训练的权重:
from keras.models import Model from keras.layers import Input, Dense, LSTM # 用Functional API重构相同结构的模型 inputs = Input(shape=(x_train_appended.shape[1], x_train_appended.shape[2])) x = LSTM(100, return_sequences=True)(inputs) x = LSTM(50, return_sequences=True)(x) x = LSTM(10)(x) outputs = Dense(1)(x) model2_func = Model(inputs=inputs, outputs=outputs) # 加载原训练好的模型权重 model2_func.set_weights(model2.get_weights()) model2_func.compile(loss='mae', optimizer='adam')
重构后再用上面的SHAP调用代码运行,警告和报错都会消失。
注意:LSTM模型的SHAP值计算速度会比普通全连接网络慢很多,不要一次性传入全量数据集做解释,分批传入或者只抽取需要解释的样本子集计算即可。
内容的提问来源于stack exchange,提问作者vmt
相关产品推荐
相关产品推荐

