R中运行predict函数遇UseMethod错误:SIS-SCAD模型预测问题
问题
在高维数据集上运行SIS(Sure Independence screening)-SCAD模型,目标是生成混淆矩阵并计算accuracy、sensitivity、specificity及F1 score等指标,但调用predict函数时出错:
model_prediction <- predict(model,X_test,type="response")
报错信息:
Error in UseMethod("predict") :
no applicable method for 'predict' applied to an object of class "c('integer', 'numeric')"
运行代码如下:
rm(list = ls()) #Installing necessary libraries library(HiDimDA) #this is where the high-dimensional data is stored library(SIS) library(caTools) data("AlonDS") set.seed(123) # Set seed for reproducibility split <- sample.split(AlonDS$grouping, SplitRatio = 0.8) train_data <- subset(AlonDS, split == TRUE) test_data <- subset(AlonDS, split == FALSE) # Extracting predictor variables and grouping variable X_train <- train_data[, -(1:6)] # Exclude non-predictor columns y_train <- train_data$grouping X_test <- test_data[, -(1:6)] # Exclude non-predictor columns y_test <- test_data$grouping # Convert grouping variable to factor y_train <- as.factor(y_train) y_test <- as.factor(y_test) # Check the structure of X_train str(X_train) # Check the structure of y_train str(y_train) # Set greedy to FALSE greedy <- FALSE # Standardize the features X_train_matrix <- scale(X_train) # Convert grouping variable to numeric (0 and 1) y_train_numeric <- as.numeric(y_train) - 1 # SIS-SCAD with the full training set model <- SIS(X_train_matrix, y_train_numeric, family = "binomial", penalty = "MCP", concavity.parameter = 3, tune = "bic", nfolds = 10, type.measure = "deviance", gamma.ebic = 1, nsis = NULL, iter = TRUE, iter.max = ifelse(greedy == FALSE, 10, floor(nrow(X_train_matrix) / log(nrow(X_train_matrix)))), varISIS = "vanilla", perm = FALSE, greedy = FALSE, greedy.size = 1, seed = NULL, standardize = TRUE) # Making predictions model_prediction <- predict(model,X_test,type="response") # Convert predicted probabilities to class labels predictions_class <- ifelse(predictions_prob > 0.5, "healthy", "colonc") # Assuming you have the true labels for comparison true_labels <- as.character(y_test) # Convert predictions_class and true_labels to factors with the same levels predictions_class <- factor(predictions_class, levels = c("healthy", "colonc")) true_labels <- factor(true_labels, levels = c("healthy", "colonc")) # Check the levels of predicted and true labels levels(predictions_class) levels(true_labels) # Create a confusion matrix conf_matrix <- confusionMatrix(predictions_class, true_labels) # Extract values from the confusion matrix conf_matrix_values <- as.table(conf_matrix) # Calculate accuracy accuracy <- sum(diag(conf_matrix_values)) / sum(conf_matrix_values) cat("Accuracy:", accuracy, "\n") # Calculate sensitivity (recall) sensitivity <- conf_matrix_values[2, 2] / sum(conf_matrix_values[2, ]) cat("Sensitivity (Recall):", sensitivity, "\n") # Calculate specificity specificity <- conf_matrix_values[1, 1] / sum(conf_matrix_values[1, ]) cat("Specificity:", specificity, "\n") # Calculate F1 score precision <- conf_matrix_values[2, 2] / sum(conf_matrix_values[, 2]) recall <- sensitivity f1_score <- 2 * (precision * recall) / (precision + recall) cat("F1 Score:", f1_score, "\n")
问题分析与解决方案
错误原因
SIS()返回值非模型对象:SIS包的SIS()函数默认返回筛选出的变量索引(整数向量),而非可用于predict()的模型对象,因此触发“无对应predict方法”的错误。- 参数缺失:未设置
return.model = TRUE,导致函数不返回训练好的模型。 - 惩罚项不符:代码中用
penalty = "MCP",但目标是SCAD模型,参数设置错误。 - 测试集未标准化:训练集做了标准化处理,测试集未同步用相同参数标准化,会导致预测偏差。
- 变量名错误:后续代码中用
predictions_prob,但实际赋值的变量是model_prediction,会引发未定义错误。 - 依赖包缺失:
confusionMatrix函数来自caret包,原代码未加载。
修正步骤
- 在
SIS()中添加return.model = TRUE,让函数返回可用于预测的模型对象。 - 将
penalty = "MCP"改为penalty = "SCAD",匹配目标模型。 - 用训练集的标准化参数(均值、标准差)处理测试集,保证数据分布一致。
- 修正变量名错误,统一使用
model_prediction。 - 加载
caret包以使用confusionMatrix函数。
修正后的完整代码
rm(list = ls()) # 加载必要库 library(HiDimDA) library(SIS) library(caTools) library(caret) # 用于生成混淆矩阵 data("AlonDS") set.seed(123) # 保证结果可复现 split <- sample.split(AlonDS$grouping, SplitRatio = 0.8) train_data <- subset(AlonDS, split == TRUE) test_data <- subset(AlonDS, split == FALSE) # 提取特征与标签 X_train <- train_data[, -(1:6)] y_train <- train_data$grouping X_test <- test_data[, -(1:6)] y_test <- test_data$grouping # 转换标签为因子 y_train <- as.factor(y_train) y_test <- as.factor(y_test) # 标准化训练集并保存参数 scale_result <- scale(X_train) X_train_matrix <- scale_result center_params <- attr(scale_result, "scaled:center") scale_params <- attr(scale_result, "scaled:scale") # 转换标签为0-1数值型 y_train_numeric <- as.numeric(y_train) - 1 # 训练SIS-SCAD模型,设置return.model=TRUE返回模型对象 model <- SIS(X_train_matrix, y_train_numeric, family = "binomial", penalty = "SCAD", # 改为SCAD惩罚项 concavity.parameter = 3, tune = "bic", nfolds = 10, type.measure = "deviance", gamma.ebic = 1, nsis = NULL, iter = TRUE, iter.max = ifelse(FALSE, 10, floor(nrow(X_train_matrix) / log(nrow(X_train_matrix)))), varISIS = "vanilla", perm = FALSE, greedy = FALSE, greedy.size = 1, seed = NULL, standardize = TRUE, return.model = TRUE) # 添加该参数返回训练好的模型 # 用训练集的标准化参数处理测试集 X_test_matrix <- scale(X_test, center = center_params, scale = scale_params) # 生成预测概率 model_prediction <- predict(model, X_test_matrix, type = "response") # 转换为类别标签 predictions_class <- ifelse(model_prediction > 0.5, "healthy", "colonc") # 统一标签的因子水平 true_labels <- as.character(y_test) predictions_class <- factor(predictions_class, levels = c("healthy", "colonc")) true_labels <- factor(true_labels, levels = c("healthy", "colonc")) # 生成混淆矩阵并输出 conf_matrix <- confusionMatrix(predictions_class, true_labels) print(conf_matrix) # 提取并打印各项指标 accuracy <- conf_matrix$overall["Accuracy"] sensitivity <- conf_matrix$byClass["Sensitivity"] specificity <- conf_matrix$byClass["Specificity"] f1_score <- conf_matrix$byClass["F1"] cat("\nAccuracy:", accuracy, "\n") cat("Sensitivity (Recall):", sensitivity, "\n") cat("Specificity:", specificity, "\n") cat("F1 Score:", f1_score, "\n")
内容的提问来源于stack exchange,提问作者Ronit
相关产品推荐
相关产品推荐

