mlr3中surv.coxph带案例权重模型预测报错的解决咨询
修复surv.coxph预测时的"variable lengths differ (found for '(weights)')"错误
问题原因
错误根源是任务中被标记为权重的列被意外移除:你执行task$select(selected_features)时,仅保留了litter、rx、sex三个特征列,把之前添加的weights列从任务数据中删除了,但task$col_roles$weight仍指向这个已不存在的列,导致预测时模型无法找到权重列,触发长度不匹配报错。
修复方案
修改任务定义代码,在select操作时保留weights列,同时确保它仅作为权重而非特征被使用:
library(mlr3) library(mlr3proba) rm(list=ls()) # Load the rats dataset and create a survival task task0 = tsk("rats") # Define task for surv.coxph learner with case weights set.seed(123) selected_features= c("litter", "rx", "sex") # 将权重列加入需保留的列列表 selected_cols = c(selected_features, "weights") task = task0$clone() task$cbind(data.frame(weights = runif(task$nrow, 1, 2))) # 选择包含权重列的所有必要列 task$select(selected_cols) task$col_roles$weight = "weights" # 将权重列从特征角色中移除,避免被当作预测特征 task$col_roles$feature = setdiff(task$col_roles$feature, "weights") task
验证修复后的完整流程
训练和预测代码无需修改,可正常执行:
# Define the learner - Cox Proportional Hazards Model (with case weights) learner = lrn("surv.coxph") # Perform a training/test split, stratified on `status` by default part = partition(task) # Train/test split balanced by status (default) # Train the learner on the training split learner$train(task, part$train) # Make predictions on the testing split p = learner$predict(task, part$test) # 现在可正常生成预测结果 p
关键说明
task$select(selected_cols)确保权重列不会被从任务数据中移除,保证模型训练和预测时都能访问到权重信息task$col_roles$feature = setdiff(...)是可选但推荐的操作,它明确将权重列排除在特征列表外,避免模型误将其当作预测变量使用
内容的提问来源于stack exchange,提问作者atg
相关产品推荐
相关产品推荐

