You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

随机森林模型执行SHAP分析时出现维度不匹配错误的解决求助

随机森林模型执行SHAP分析时出现维度不匹配错误的解决求助

我现在在运行随机森林模型,想要获取特征重要性,于是尝试做SHAP分析,但每次尝试绘制SHAP值的时候,都会遇到这个错误:

DimensionError: Length of features is not equal to the length of shap_values.

我完全搞不懂哪里出问题了——用同一个数据集跑XGBoost模型的时候一切正常,SHAP图能正常显示,但换成随机森林就不行了。这是个二分类任务,数据集已经预处理过了,列要么是0-1的二值列要么是数值列。

以下是我的Python代码:

from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, confusion_matrix

# Remove the primary key column 'id' from the features

features = result.drop(columns=['PQ2', 'id'])  # Drop target and ID columns
target = result['PQ2']  # Target variable

# Split data into training and testing sets with 80-20 ratio
X_train, X_test, y_train, y_test = train_test_split(features, target, test_size=0.2, random_state=42)
 
# Initialize Random Forest classifier
rf_model = RandomForestClassifier(n_estimators=100, random_state=42)

# Fit the model on the training data
rf_model.fit(X_train, y_train)

# Make predictions
y_pred = rf_model.predict(X_test)

import shap

# Create a Tree SHAP explainer for the Random Forest model
explainer = shap.TreeExplainer(rf_model)

# Calculate SHAP values for the test set
shap_values = explainer.shap_values(X_test)

# Plot a SHAP summary plot
shap.summary_plot(shap_values, X_test, feature_names=features_names)

# Plot a SHAP bar plot for global feature importance

shap.summary_plot(shap_values, X_test, feature_names=features_names, plot_type="bar")

我检查了维度,测试集的形状是(829,22),但随机森林输出的SHAP值形状却是(22,2),这完全不对,我实在不知道该怎么修复这个问题。

备注:内容来源于stack exchange,提问作者Starterkit07

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 19:09:31