mlr3shiny中二类分类PipeOp学习器的概率值获取及DALEX集成问题
我太懂你这种头疼的感觉了——好不容易把mlr3和Shiny搭起来,结果因为PipeOpThreshold的限制,拿不到DALEX需要的概率值,简直卡得死死的。咱们先把问题拆透,再一步步解决:
问题根源
你遇到的PipeOpError: $predict_type for PipeOpThreshold is read-only是因为PipeOpThreshold的设计就是专门输出**response类型(类别预测)**的,一旦把它加到你的graph里,整个GraphLearner的预测类型就会被锁死成response,自然输出不了概率矩阵,而DALEX做模型解释恰恰需要概率值来计算各种指标。
解决方案:两种思路任你选
思路一:移除PipeOpThreshold,改用Predict时指定阈值(推荐)
既然PipeOpThreshold是麻烦的源头,咱们不用它就行——mlr3本身支持在predict阶段直接指定阈值来生成类别预测,同时保留概率输出。
修改你的createGraphLearner函数:
createGraphLearner <- function(selectedlearner) { # 统一设置所有学习器的predict_type为prob,二类分类也不例外 learner <- lrn(input[[selectedlearner]], predict_type = "prob") if(input[["Task_robustify"]]){ graph <- pipeline_robustify(currenttask$task, learner) %>>% learner } else { graph <- as_graph(po("learner", learner)) } plot(graph) # 删掉添加PipeOpThreshold的代码,后续手动处理阈值 # if (isTRUE(currenttask$task$properties == "twoclass")) graph <- graph %>>% po("threshold") return(as_learner(graph)) }
这样训练出来的GraphLearner始终保持predict_type="prob",既能输出概率给DALEX,又能在需要类别预测时,通过predict函数的threshold参数生成:
# 假设trained_learner是训练好的模型 pred <- trained_learner$predict(currenttask$task, threshold = 0.5) # 可根据需求调整阈值 # pred里同时包含response(类别)和prob(概率矩阵)
思路二:保留PipeOpThreshold,让Graph同时输出概率和类别
如果你一定要保留PipeOpThreshold在流程里,可以通过数据流分支的方式,同时保留概率输出和类别输出:
createGraphLearner <- function(selectedlearner) { if (!isTRUE(currenttask$task$properties == "twoclass")) { learner <- lrn(input[[selectedlearner]]) } else { learner <- lrn(input[[selectedlearner]], predict_type = "prob") } if(input[["Task_robustify"]]){ graph <- pipeline_robustify(currenttask$task, learner) %>>% learner } else { graph <- as_graph(po("learner", learner)) } plot(graph) if (isTRUE(currenttask$task$properties == "twoclass")) { # 复制数据流:一份走阈值生成类别,一份保留原始概率 graph <- graph %>>% po("copy", outnum = 2) %>>% gunion(list( po("threshold"), # 生成response类别 po("nop") # 啥也不做,保留prob )) %>>% po("cbind") # 合并两个分支的输出 } return(as_learner(graph)) }
这种方式下,预测结果会同时包含类别和概率,但需要注意DALEX对这种复合输出的兼容性,所以还是思路一更稳妥。
对接DALEX的正确姿势
用DALEXtra包的explain_mlr3函数直接对接训练好的GraphLearner就行,因为我们已经确保它能输出概率:
library(DALEXtra) # 生成解释器 explainer <- explain_mlr3( model = trained_learner, data = currenttask$task$data(), y = currenttask$task$truth() ) # 接下来就可以正常做DALEX的各种分析了,比如SHAP值、局部解释
这样就能完美解决你的问题啦——既不影响Shiny里的分类预测,又能给DALEX提供需要的概率数据。
备注:内容来源于stack exchange,提问作者DominicHOST

