You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R使用automl建模调用caret::confusionMatrix报levels数量不匹配错误

报错原因分析

你遇到的confusionMatrix.default(test$Y, prediction) : the data cannot have more levels than the reference报错,核心是caret包的混淆矩阵函数要求:作为参考值的真实标签(第一个输入参数)的因子水平,必须完全覆盖预测值(第二个输入参数)的所有因子水平。
你代码存在两个问题直接导致报错:

  • 语法错误:automl_train_manual函数调用时,subset(train, select = -c(Y)后面漏写了右括号,先触发语法层面的运行异常
  • 因子水平不匹配:你将预测结果转为因子时没有指定水平,如果模型预测结果全部为类别2,生成的prediction因子就只有"2"一个水平,和真实标签test$Y的"1"、"2"两个水平不匹配,就会触发该报错。另外你写的预测值判断逻辑冗余,两层ifelse其实等价于大于1.5返回2、否则返回1。
修正后可直接运行的完整代码
library(automl)
library(rsample)
library(dplyr)

# 直接使用你提供的数据集结构
itog=structure(list(r = c(408L, 450L, 667L, 477L, 374L, 260L, 419L, 
441L, 658L, 374L, 333L, 313L, 404L, 432L, 458L, 457L, 286L, 286L, 
259L, 238L, 230L, 214L, 259L, 201L, 232L, 235L, 233L, 252L, 271L, 
259L, 235L, 210L, 206L, 244L, 218L, 211L, 246L, 242L, 217L, 255L, 
277L, 262L, 280L, 278L, 289L, 271L, 236L, 249L, 235L, 248L, 263L, 
217L, 263L, 300L, 242L, 269L, 280L, 301L, 317L, 247L, 218L, 209L, 
237L, 253L), g = c(384L, 418L, 656L, 480L, 397L, 341L, 441L, 
461L, 597L, 422L, 394L, 314L, 433L, 451L, 464L, 406L, 260L, 262L, 
216L, 244L, 230L, 222L, 234L, 242L, 211L, 232L, 231L, 235L, 281L, 
251L, 241L, 194L, 192L, 222L, 234L, 269L, 220L, 221L, 226L, 224L, 
210L, 274L, 272L, 281L, 247L, 264L, 226L, 247L, 254L, 223L, 255L, 
217L, 238L, 260L, 248L, 247L, 272L, 313L, 323L, 254L, 202L, 202L, 
297L, 269L), b = c(372L, 477L, 617L, 495L, 414L, 314L, 430L, 
487L, 623L, 514L, 351L, 343L, 422L, 433L, 449L, 425L, 363L, 326L, 
312L, 308L, 286L, 259L, 243L, 265L, 280L, 298L, 303L, 301L, 359L, 
298L, 295L, 265L, 263L, 298L, 275L, 332L, 303L, 287L, 305L, 285L, 
298L, 339L, 339L, 359L, 300L, 293L, 289L, 277L, 326L, 312L, 311L, 
289L, 317L, 325L, 343L, 332L, 379L, 422L, 415L, 326L, 292L, 257L, 
288L, 306L), Y = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 
1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 
2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 
2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 
2L, 2L, 2L, 2L, 2L)), class = "data.frame", row.names = c(NA, 
-64L))

itog$Y=as.factor(itog$Y)
set.seed(123)
indexes = createDataPartition(itog$Y, p = .7, list = F)
train = itog[indexes, ]
test = itog[-indexes, ]

# 补全之前漏写的右括号
amlmodel = automl_train_manual(Xref = subset(train, select = -c(Y)),
                               Yref = subset(train, select = c(Y))$Y %>% as.numeric(),
                               hpar = list(learningrate = 0.01,
                                           minibatchsize = 2^2,
                                           numiterations = 60))
prediction = automl_predict(model = amlmodel, X = test[,1:3]) 
# 简化冗余的判断逻辑
prediction = ifelse(prediction > 1.5, 2, 1) 
# 转因子时强制指定水平和真实标签一致,避免水平缺失
prediction = factor(prediction, levels = levels(test$Y))
caret::confusionMatrix(test$Y, prediction)
异常排查方法

如果调整后仍然出现报错,可以先运行以下代码验证两个变量的水平是否匹配:

print(levels(test$Y))
print(levels(prediction))

如果两者水平不一致,说明你的模型训练效果较差,只预测出了部分类别,可以适当调大迭代次数、调整学习率优化模型效果。

内容的提问来源于stack exchange,提问作者psysky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 15:45:06