基于全哑变量的BART分类建模问题咨询
问题背景
我尝试在预测变量和响应变量均为哑变量的场景下,应用BART进行分类建模。哑变量由取值范围为-4到4的分类变量转换而来:负值设为0,正值设为1,同时我也保留了原始分类版本的数据。
我的预测变量是一个648×48的0-1哑变量矩阵,其中缺失值(NA)占比高达70%;响应变量无缺失值,共648个样本。我在RStudio中用R语言建模,运行代码后结果不尽人意:
bart_machine = build_bart_machine(predictors, response_var,use_missing_data = TRUE, use_missing_data_dummies_as_covars = TRUE) bart_machine$confusion_matrix
运行后得到的混淆矩阵为NULL,同时输出信息如下:
bartMachine v1.3.4.1 for regression Missing data feature ON training data size: n = 638 and p = 96 built in 5.8 secs on 8 cores, 50 trees, 250 burn-in and 1000 post. samples sigsq est for y beforehand: 0.016 avg sigsq estimate after burn-in: 0.00314 in-sample statistics: L1 = 9.91 L2 = 1.01 rmse = 0.04 Pseudo-Rsq = 0.9547 p-val for shapiro-wilk test of normality of residuals: 0 p-val for zero-mean noise: 0.99451
我有三个问题:
- 70%的缺失值是否会导致无法正常进行BART建模?
- 当前设置下是否应该能生成混淆矩阵?
- 使用原始分类版本的数据是否更有帮助,还是当前不佳结果源于更核心的问题?
补充说明:我已进行预测,但效果极差:
# Extract the feature names feature_names <- bart_machine[["training_data_features_with_missing_features"]] # Remove the "M_" prefix feature_names <- gsub("^M_", "", feature_names) # Update the bart_machine object with the renamed feature names bart_machine[["training_data_features_with_missing_features"]] <- feature_names # Inspect feature names in bart_machine training data training_data <- bart_machine[["model_matrix_training_data"]] training_feature_names <- colnames(training_data) # Remove the "M_" prefix from the column names training_feature_names <- gsub("^M_", "", training_feature_names) # Update the bart_machine object with the renamed columns colnames(training_data) <- training_feature_names bart_machine[["model_matrix_training_data"]] <- training_data # Impute missing data in albany2005_predictors using the same method imputed_data_2005 <- mice(albany2005[, -1], m = 5, method = 'pmm', maxit = 50, seed = 500) complete_data_2005 <- complete(imputed_data_2005, 1) # Remove the "M_" prefix from the column names in albany2005_predictors colnames(complete_data_2005) <- gsub("^M_", "", colnames(complete_data_2005)) # Ensure albany2005_predictors has the same columns as the training data, excluding 'y_remaining' required_cols <- setdiff(training_feature_names, "y_remaining") albany2005_predictors <- complete_data_2005[, required_cols, drop = FALSE] # Create a logical vector for non-NA values in albany2005_response non_na_indices <- !is.na(albany2005_response) # Subset predicted_values and albany2005_response using non_na_indices non_na_predicted_values <- predicted_values[non_na_indices] non_na_actual_values <- albany2005_response[non_na_indices] # Predict using the bartMachine model on the 2005 data predicted_values <- predict(bart_machine, albany2005_predictors, type = "class") # Calculate RMSE # Convert predicted values to numeric non_na_predicted_values <- as.numeric(as.character(non_na_predicted_values)) # Convert actual values to numeric non_na_actual_values <- as.numeric(as.character(non_na_actual_values)) rmse <- sqrt(mean((non_na_predicted_values - non_na_actual_values)^2)) # Print the RMSE print(paste("RMSE: ", rmse)) ###very high RMSE!!!
解答
1. 70%缺失值对BART建模的影响
70%的高缺失值会给BART建模带来严重阻碍:
- bartMachine的缺失值处理机制是生成缺失指示变量,但单个特征缺失率极高时,这些指示变量几乎没有区分度,模型无法学到有效信息。
- 高缺失率会大幅降低数据信息密度,即使模型能运行,也极易出现训练集过拟合(你看到的训练集RMSE极低但预测RMSE极高就是典型表现),泛化能力极差。
2. 混淆矩阵为NULL的原因
从输出信息看,你当前训练的是回归模型(第一行明确标注bartMachine v1.3.4.1 for regression),而混淆矩阵是分类模型的专属输出。问题出在响应变量类型:
- 如果响应变量是0-1数值型而非因子型,bartMachine默认拟合回归模型,不会生成混淆矩阵。
- 解决方法:将响应变量转换为因子型再训练,执行
response_var <- as.factor(response_var),此时bartMachine会自动拟合分类模型,训练完成后bart_machine$confusion_matrix就会有结果。
3. 原始分类数据是否更优?
使用原始分类版本的数据(-4到4的分类变量)大概率会比当前的0-1哑变量更好:
- 你把多分类变量强行二值化,丢失了原始变量的层级信息(比如-4和-1都被设为0,原始差异被抹除),大幅降低了特征的预测能力。
- 但核心问题还是高缺失率,即使换用原始数据,70%的缺失值依然会让模型难以学到有效模式。建议优先处理缺失值:
- 先分析缺失机制是随机还是非随机,若为非随机需先排查缺失原因;
- 尝试更可靠的缺失值插补方法(比如用bartMachine自带的插补而非mice,保持模型前后处理一致);
- 考虑删除缺失率过高的特征(比如缺失率超过50%的特征直接丢弃),减少噪声特征对模型的干扰。
另外,你后续预测时直接修改bartMachine对象内部的特征名称是风险操作,可能导致预测时特征匹配错误,这也可能是预测效果极差的原因之一,建议不要修改模型对象的内部属性,而是在训练前就统一处理好特征名称。
内容的提问来源于stack exchange,提问作者Lusian
相关产品推荐
相关产品推荐

