使用IML计算Shapley值时results与fit为空的问题求助
问题原因与解决步骤
核心问题
iml::Shapley$new()初始化对象时不会自动计算Shapley值,同时party::cforest模型与iml交互时需要明确预测类型,这两个点导致你得到的shapley_vals对象中results和fit为NULL。
具体修正方案
1. 适配cforest的预测函数
cforest的默认预测输出格式需要调整,确保iml能正确获取回归结果:
# 自定义预测函数,指定回归任务的输出类型 predict_fun <- function(model, newdata) { predict(model, newdata = newdata, type = "response") }
2. 创建正确的Predictor对象
用自定义的预测函数初始化Predictor:
predictor <- iml::Predictor$new( model = cvforest.model.train, data = X, y = y, predict.function = predict_fun )
3. 指定样本并计算Shapley值
必须指定要解释的样本(单样本或批量样本)才会触发计算:
- 单样本计算:
# 对训练集第1个样本计算Shapley值 shapley_single <- iml::Shapley$new(predictor, x.interest = X[1, ]) # 查看结果 print(shapley_single$results)
- 批量样本计算:
# 随机抽取50个样本计算Shapley值 shapley_vals <- iml::Shapley$new(predictor, sample.size = 50) # 查看结果 str(shapley_vals$results)
完整修正代码
# training random forest model set.seed(8431) cvforest.train <- train(x = train[, c(3:122, 124:138)], y = train[, 123], method = "cforest", metric = "RMSE", trControl = trainControl(method = "cv", number = 5), controls = cforest_unbiased(ntree = 1000, minsplit = 5, minbucket = 5) ) cvforest.model.train <- cvforest.train$finalModel #final model object # calculating shapley values X <- train[, c(3:122, 124:138)] y <- train[, 123] # 自定义适配cforest的预测函数 predict_fun <- function(model, newdata) { predict(model, newdata = newdata, type = "response") } # 创建Predictor对象 predictor <- iml::Predictor$new( model = cvforest.model.train, data = X, y = y, predict.function = predict_fun ) # 计算批量样本的Shapley值 shapley_vals <- iml::Shapley$new(predictor, sample.size = 50) # 查看结果 str(shapley_vals$results)
内容的提问来源于stack exchange,提问作者Sarah Vogel
相关产品推荐
相关产品推荐

