You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于全哑变量的BART分类建模问题咨询

问题背景

我尝试在预测变量和响应变量均为哑变量的场景下,应用BART进行分类建模。哑变量由取值范围为-4到4的分类变量转换而来:负值设为0,正值设为1,同时我也保留了原始分类版本的数据。

我的预测变量是一个648×48的0-1哑变量矩阵,其中缺失值(NA)占比高达70%;响应变量无缺失值,共648个样本。我在RStudio中用R语言建模,运行代码后结果不尽人意:

bart_machine = build_bart_machine(predictors, response_var,use_missing_data = TRUE, use_missing_data_dummies_as_covars = TRUE)
bart_machine$confusion_matrix

运行后得到的混淆矩阵为NULL,同时输出信息如下:

bartMachine v1.3.4.1 for regression

Missing data feature ON
training data size: n = 638 and p = 96 
built in 5.8 secs on 8 cores, 50 trees, 250 burn-in and 1000 post. samples

sigsq est for y beforehand: 0.016 
avg sigsq estimate after burn-in: 0.00314 

in-sample statistics:
 L1 = 9.91 
 L2 = 1.01 
 rmse = 0.04 
 Pseudo-Rsq = 0.9547
p-val for shapiro-wilk test of normality of residuals: 0 
p-val for zero-mean noise: 0.99451 

我有三个问题:

  1. 70%的缺失值是否会导致无法正常进行BART建模?
  2. 当前设置下是否应该能生成混淆矩阵?
  3. 使用原始分类版本的数据是否更有帮助,还是当前不佳结果源于更核心的问题?

补充说明:我已进行预测,但效果极差:

# Extract the feature names
feature_names <- bart_machine[["training_data_features_with_missing_features"]]

# Remove the "M_" prefix
feature_names <- gsub("^M_", "", feature_names)

# Update the bart_machine object with the renamed feature names
bart_machine[["training_data_features_with_missing_features"]] <- feature_names


# Inspect feature names in bart_machine training data
training_data <- bart_machine[["model_matrix_training_data"]]
training_feature_names <- colnames(training_data)

# Remove the "M_" prefix from the column names
training_feature_names <- gsub("^M_", "", training_feature_names)

# Update the bart_machine object with the renamed columns
colnames(training_data) <- training_feature_names
bart_machine[["model_matrix_training_data"]] <- training_data

# Impute missing data in albany2005_predictors using the same method
imputed_data_2005 <- mice(albany2005[, -1], m = 5, method = 'pmm', maxit = 50, seed = 500)
complete_data_2005 <- complete(imputed_data_2005, 1)

# Remove the "M_" prefix from the column names in albany2005_predictors
colnames(complete_data_2005) <- gsub("^M_", "", colnames(complete_data_2005))

# Ensure albany2005_predictors has the same columns as the training data, excluding 'y_remaining'
required_cols <- setdiff(training_feature_names, "y_remaining")
albany2005_predictors <- complete_data_2005[, required_cols, drop = FALSE]



# Create a logical vector for non-NA values in albany2005_response
non_na_indices <- !is.na(albany2005_response)

# Subset predicted_values and albany2005_response using non_na_indices
non_na_predicted_values <- predicted_values[non_na_indices]
non_na_actual_values <- albany2005_response[non_na_indices]


# Predict using the bartMachine model on the 2005 data
predicted_values <- predict(bart_machine, albany2005_predictors, type = "class")

# Calculate RMSE
# Convert predicted values to numeric
non_na_predicted_values <- as.numeric(as.character(non_na_predicted_values))

# Convert actual values to numeric
non_na_actual_values <- as.numeric(as.character(non_na_actual_values))

rmse <- sqrt(mean((non_na_predicted_values - non_na_actual_values)^2))

# Print the RMSE
print(paste("RMSE: ", rmse)) ###very high RMSE!!!

解答

1. 70%缺失值对BART建模的影响

70%的高缺失值会给BART建模带来严重阻碍:

  • bartMachine的缺失值处理机制是生成缺失指示变量,但单个特征缺失率极高时,这些指示变量几乎没有区分度,模型无法学到有效信息。
  • 高缺失率会大幅降低数据信息密度,即使模型能运行,也极易出现训练集过拟合(你看到的训练集RMSE极低但预测RMSE极高就是典型表现),泛化能力极差。

2. 混淆矩阵为NULL的原因

从输出信息看,你当前训练的是回归模型(第一行明确标注bartMachine v1.3.4.1 for regression),而混淆矩阵是分类模型的专属输出。问题出在响应变量类型:

  • 如果响应变量是0-1数值型而非因子型,bartMachine默认拟合回归模型,不会生成混淆矩阵。
  • 解决方法:将响应变量转换为因子型再训练,执行response_var <- as.factor(response_var),此时bartMachine会自动拟合分类模型,训练完成后bart_machine$confusion_matrix就会有结果。

3. 原始分类数据是否更优?

使用原始分类版本的数据(-4到4的分类变量)大概率会比当前的0-1哑变量更好:

  • 你把多分类变量强行二值化,丢失了原始变量的层级信息(比如-4和-1都被设为0,原始差异被抹除),大幅降低了特征的预测能力。
  • 但核心问题还是高缺失率,即使换用原始数据,70%的缺失值依然会让模型难以学到有效模式。建议优先处理缺失值:
    • 先分析缺失机制是随机还是非随机,若为非随机需先排查缺失原因;
    • 尝试更可靠的缺失值插补方法(比如用bartMachine自带的插补而非mice,保持模型前后处理一致);
    • 考虑删除缺失率过高的特征(比如缺失率超过50%的特征直接丢弃),减少噪声特征对模型的干扰。

另外,你后续预测时直接修改bartMachine对象内部的特征名称是风险操作,可能导致预测时特征匹配错误,这也可能是预测效果极差的原因之一,建议不要修改模型对象的内部属性,而是在训练前就统一处理好特征名称。


内容的提问来源于stack exchange,提问作者Lusian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 04:57:32