R使用automl建模调用caret::confusionMatrix报levels数量不匹配错误
报错原因分析
你遇到的confusionMatrix.default(test$Y, prediction) : the data cannot have more levels than the reference报错,核心是caret包的混淆矩阵函数要求:作为参考值的真实标签(第一个输入参数)的因子水平,必须完全覆盖预测值(第二个输入参数)的所有因子水平。
你代码存在两个问题直接导致报错:
- 语法错误:
automl_train_manual函数调用时,subset(train, select = -c(Y)后面漏写了右括号,先触发语法层面的运行异常 - 因子水平不匹配:你将预测结果转为因子时没有指定水平,如果模型预测结果全部为类别2,生成的
prediction因子就只有"2"一个水平,和真实标签test$Y的"1"、"2"两个水平不匹配,就会触发该报错。另外你写的预测值判断逻辑冗余,两层ifelse其实等价于大于1.5返回2、否则返回1。
修正后可直接运行的完整代码
library(automl) library(rsample) library(dplyr) # 直接使用你提供的数据集结构 itog=structure(list(r = c(408L, 450L, 667L, 477L, 374L, 260L, 419L, 441L, 658L, 374L, 333L, 313L, 404L, 432L, 458L, 457L, 286L, 286L, 259L, 238L, 230L, 214L, 259L, 201L, 232L, 235L, 233L, 252L, 271L, 259L, 235L, 210L, 206L, 244L, 218L, 211L, 246L, 242L, 217L, 255L, 277L, 262L, 280L, 278L, 289L, 271L, 236L, 249L, 235L, 248L, 263L, 217L, 263L, 300L, 242L, 269L, 280L, 301L, 317L, 247L, 218L, 209L, 237L, 253L), g = c(384L, 418L, 656L, 480L, 397L, 341L, 441L, 461L, 597L, 422L, 394L, 314L, 433L, 451L, 464L, 406L, 260L, 262L, 216L, 244L, 230L, 222L, 234L, 242L, 211L, 232L, 231L, 235L, 281L, 251L, 241L, 194L, 192L, 222L, 234L, 269L, 220L, 221L, 226L, 224L, 210L, 274L, 272L, 281L, 247L, 264L, 226L, 247L, 254L, 223L, 255L, 217L, 238L, 260L, 248L, 247L, 272L, 313L, 323L, 254L, 202L, 202L, 297L, 269L), b = c(372L, 477L, 617L, 495L, 414L, 314L, 430L, 487L, 623L, 514L, 351L, 343L, 422L, 433L, 449L, 425L, 363L, 326L, 312L, 308L, 286L, 259L, 243L, 265L, 280L, 298L, 303L, 301L, 359L, 298L, 295L, 265L, 263L, 298L, 275L, 332L, 303L, 287L, 305L, 285L, 298L, 339L, 339L, 359L, 300L, 293L, 289L, 277L, 326L, 312L, 311L, 289L, 317L, 325L, 343L, 332L, 379L, 422L, 415L, 326L, 292L, 257L, 288L, 306L), Y = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L)), class = "data.frame", row.names = c(NA, -64L)) itog$Y=as.factor(itog$Y) set.seed(123) indexes = createDataPartition(itog$Y, p = .7, list = F) train = itog[indexes, ] test = itog[-indexes, ] # 补全之前漏写的右括号 amlmodel = automl_train_manual(Xref = subset(train, select = -c(Y)), Yref = subset(train, select = c(Y))$Y %>% as.numeric(), hpar = list(learningrate = 0.01, minibatchsize = 2^2, numiterations = 60)) prediction = automl_predict(model = amlmodel, X = test[,1:3]) # 简化冗余的判断逻辑 prediction = ifelse(prediction > 1.5, 2, 1) # 转因子时强制指定水平和真实标签一致,避免水平缺失 prediction = factor(prediction, levels = levels(test$Y)) caret::confusionMatrix(test$Y, prediction)
异常排查方法
如果调整后仍然出现报错,可以先运行以下代码验证两个变量的水平是否匹配:
print(levels(test$Y)) print(levels(prediction))
如果两者水平不一致,说明你的模型训练效果较差,只预测出了部分类别,可以适当调大迭代次数、调整学习率优化模型效果。
内容的提问来源于stack exchange,提问作者psysky
相关产品推荐
相关产品推荐

