随机森林模型执行SHAP分析时出现维度不匹配错误的解决求助
随机森林模型执行SHAP分析时出现维度不匹配错误的解决求助
我现在在运行随机森林模型,想要获取特征重要性,于是尝试做SHAP分析,但每次尝试绘制SHAP值的时候,都会遇到这个错误:
DimensionError: Length of features is not equal to the length of shap_values.
我完全搞不懂哪里出问题了——用同一个数据集跑XGBoost模型的时候一切正常,SHAP图能正常显示,但换成随机森林就不行了。这是个二分类任务,数据集已经预处理过了,列要么是0-1的二值列要么是数值列。
以下是我的Python代码:
from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import train_test_split, cross_val_score from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, confusion_matrix # Remove the primary key column 'id' from the features features = result.drop(columns=['PQ2', 'id']) # Drop target and ID columns target = result['PQ2'] # Target variable # Split data into training and testing sets with 80-20 ratio X_train, X_test, y_train, y_test = train_test_split(features, target, test_size=0.2, random_state=42) # Initialize Random Forest classifier rf_model = RandomForestClassifier(n_estimators=100, random_state=42) # Fit the model on the training data rf_model.fit(X_train, y_train) # Make predictions y_pred = rf_model.predict(X_test) import shap # Create a Tree SHAP explainer for the Random Forest model explainer = shap.TreeExplainer(rf_model) # Calculate SHAP values for the test set shap_values = explainer.shap_values(X_test) # Plot a SHAP summary plot shap.summary_plot(shap_values, X_test, feature_names=features_names) # Plot a SHAP bar plot for global feature importance shap.summary_plot(shap_values, X_test, feature_names=features_names, plot_type="bar")
我检查了维度,测试集的形状是(829,22),但随机森林输出的SHAP值形状却是(22,2),这完全不对,我实在不知道该怎么修复这个问题。
备注:内容来源于stack exchange,提问作者Starterkit07
相关产品推荐
相关产品推荐

