You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

shap.TreeExplainer与shap.Explainer条形图差异及特征重要性选型咨询

两种SHAP条形图的差异及特征重要性分析选择

问题背景

我基于1000个样本的训练集与500个样本的测试集训练随机森林分类器,生成SHAP值条形图时采用两种方式,得到的结果存在差异,想明确两者的差异点,以及哪种更适合用于特征重要性分析:

方式1:使用shap.TreeExplainer处理训练集

shap_values_Tree_tr = shap.TreeExplainer(clf.best_estimator_).shap_values(X_train)
shap.summary_plot(shap_values_Tree_tr, X_train)  # 若需条形图需添加`plot_type="bar"`参数

方式2:使用shap.Explainer处理测试集

explainer2 = shap.Explainer(clf.best_estimator_.predict, X_test)
shap_values = explainer2(X_test)
shap.plots.bar(shap_values)

完整代码如下:

from sklearn.datasets import make_classification
import seaborn as sns
import numpy as np
import pandas as pd
from matplotlib import pyplot as plt
import pickle
import joblib
import warnings
import shap
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import RandomizedSearchCV, GridSearchCV

f, (ax1,ax2) = plt.subplots(nrows=1, ncols=2,figsize=(20,8))
# Generate noisy Data
X_train,y_train = make_classification(n_samples=1000, 
                          n_features=50, 
                          n_informative=9, 
                          n_redundant=0, 
                          n_repeated=0, 
                          n_classes=10, 
                          n_clusters_per_class=1,
                          class_sep=9,
                          flip_y=0.2,
                          #weights=[0.5,0.5], 
                          random_state=17)

X_test,y_test = make_classification(n_samples=500, 
                          n_features=50, 
                          n_informative=9, 
                          n_redundant=0, 
                          n_repeated=0, 
                          n_classes=10, 
                          n_clusters_per_class=1,
                          class_sep=9,
                          flip_y=0.2,
                          #weights=[0.5,0.5], 
                          random_state=17)

model = RandomForestClassifier()

parameter_space = {
    'n_estimators': [10,50,100],
    'criterion': ['gini', 'entropy'],
    'max_depth': np.linspace(10,50,11),
}

clf = GridSearchCV(model, parameter_space, cv = 5, scoring = "accuracy", verbose = True) # model
my_model = clf.fit(X_train,y_train)
print(f'Best Parameters: {clf.best_params_}')

# save the model to disk
filename = f'Testt-RF.sav'
pickle.dump(clf, open(filename, 'wb'))

shap_values_Tree_tr = shap.TreeExplainer(clf.best_estimator_).shap_values(X_train)
shap.summary_plot(shap_values_Tree_tr, X_train)

explainer2 = shap.Explainer(clf.best_estimator_.predict, X_test)
shap_values = explainer2(X_test)

shap.plots.bar(shap_values)

两种方式的核心差异

1. 解释器计算原理不同

  • shap.TreeExplainer是树模型专属解释器,利用树结构的分裂规则精确计算SHAP值,计算速度快、无近似误差,完全适配随机森林这类树集成模型。
  • shap.Explainer是通用解释器(实际为shap.PermutationExplainer的别名),通过特征置换的方式近似估算SHAP值,适用于所有类型模型,但计算成本高、结果存在一定误差。

2. 分析的数据集分布不同

  • 方式1针对训练集计算,反映的是特征在模型训练数据中的贡献程度,贴合模型的学习逻辑。
  • 方式2针对测试集计算,反映的是特征在新数据上对模型预测的贡献程度;测试集与训练集虽生成参数一致,但样本存在差异(如flip_y引入的噪声),特征贡献可能因数据分布细微变化而不同。

3. 可视化的统计逻辑差异

  • 若方式1使用plot_type="bar"生成条形图,展示的是训练集每个特征SHAP值绝对值的均值,体现特征对模型输出的平均影响强度。
  • 方式2的shap.plots.bar展示的是测试集每个特征SHAP值绝对值的均值,体现特征对测试样本预测的平均影响强度。

特征重要性分析的选择建议

  1. 优先选择shap.TreeExplainer+目标数据集:
    对于随机森林这类树模型,TreeExplainer的精确性和效率远高于通用解释器。如果要分析测试集的特征重要性,建议直接用TreeExplainer处理测试集,代码示例:

    explainer_tree_test = shap.TreeExplainer(clf.best_estimator_)
    shap_values_tree_test = explainer_tree_test.shap_values(X_test)
    shap.summary_plot(shap_values_tree_test, X_test, plot_type="bar")
    
  2. 场景化选择:

    • 若要验证模型在训练阶段学到的特征优先级,选TreeExplainer+训练集,结果更准确,能反映模型的核心依赖特征。
    • 若要检查模型在新数据上的特征表现(比如排查数据漂移、验证部署效果),选针对测试集的分析,但优先用TreeExplainer而非通用Explainer。

内容的提问来源于stack exchange,提问作者Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 05:36:18