使用vivid包计算特征重要性时出错的原因及解决方法
问题:vivid包vivi函数运行报错的原因与修复方案
场景与数据
在R中使用vivid包的vivi函数探索变量特征重要性,现有数据结构如下:
> str(data_s) tibble [100 × 3] (S3: tbl_df/tbl/data.frame) $ y : num [1:100] 0 0 0 0 0 0 0 0 1 0 ... $ x1: num [1:100] 34 34 35 35 35 34 6 34 34 6 ... $ x2: num [1:100] 1 1 4 0 3 4 5 2 4 1 ... - attr(*, "na.action")= 'omit' Named int [1:197659] 4 5 6 7 9 14 19 20 24 27 ... ..- attr(*, "names")= chr [1:197659] "4" "5" "6" "7" ...
用户编写的代码
library("vivid") library("dplyr") library("xgboost") y=data_s["y"] x=data_s[,c("x1","x2")] gbst <- xgboost(data = as.matrix(x), label = as.matrix(y), nrounds = 600) pFun <- function(fit, data, ...) predict(fit, as.matrix(x)) viviGBst <- vivi(fit = gbst, data = data_s, response = "y", reorder = FALSE, normalized = FALSE, predictFun = pFun)
报错信息
Error: ! Assigned data `predict(x, data = X[, cols, drop = FALSE])` must be compatible with existing data. ✖ Existing data has 5000 rows. ✖ Assigned data has 100 rows. ℹ Only vectors of size 1 are recycled. Run `rlang::last_error()` to see where the error occurred.
报错原因
核心问题在于自定义的pFun函数硬编码使用了全局变量x(仅100行),而vivi函数内部默认会生成5000行的模拟数据集(由n.grid参数控制,默认值为5000)来计算变量重要性。当vivi调用pFun时,会将内部生成的5000行数据传入data参数,但pFun未使用该参数,反而一直调用原来的100行x,导致预测结果行数与vivi内部数据集行数不匹配,触发报错。
修复方案
修改pFun函数,使用传入的data参数构建预测用矩阵,而非全局变量x。具体代码如下:
# 修正后的预测函数 pFun <- function(fit, data, ...) { # 从传入的data中提取特征列(x1、x2),转为矩阵后预测 predict(fit, as.matrix(data[, c("x1", "x2")])) } # 重新运行vivi viviGBst <- vivi(fit = gbst, data = data_s, response = "y", reorder = FALSE, normalized = FALSE, predictFun = pFun)
补充说明
如果特征列不止x1和x2,可以更通用地提取除响应变量外的所有列:
pFun <- function(fit, data, ...) { # 移除响应变量列,剩余列为特征 features <- data[, setdiff(names(data), "y")] predict(fit, as.matrix(features)) }
这样pFun就能适配vivi内部生成的任意行数的数据集,保证预测结果行数与内部数据集一致。
内容的提问来源于stack exchange,提问作者oercim
相关产品推荐
相关产品推荐

