You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XGBoost生存模型训练报错:标签DataFrame含多列如何解决?

问题描述

我正在开发XGBoost Survival模型,代码片段如下:

X = df_High_School[['Gender', 'Lived_both_Parents', 'Moth_Born_in_Canada', 'Father_Born_in_Canada','Born_in_Canada','Aboriginal','Visible_Minority']]  # covariates 
y = df_High_School[['time_to_event', 'event']]  # time to event and event indicator

#split the data into training and test sets 
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)

#Develop the model 
model = xgb.XGBRegressor(objective='survival:cox')

运行model.fit(X_train, y_train)时出现如下错误:

ValueError                                Traceback (most recent call last)
<ipython-input-9-1c5a15fa4b2b> in <module>
     18 
     19 # fit the model to the training data
---> 20 model.fit(X_train, y_train)
     21 
     22 # make predictions on the test set

2 frames
/usr/local/lib/python3.8/dist-packages/xgboost/core.py in _maybe_pandas_label(label)
    261     if isinstance(label, DataFrame):
    262         if len(label.columns) > 1:
---> 263             raise ValueError('DataFrame for label cannot have multiple columns')
    264 
    265         label_dtypes = label.dtypes

ValueError: DataFrame for label cannot have multiple columns

由于这是生存模型,需要time_to_event和event两列分别表示事件时间和事件指示器,尝试将DataFrame转为Numpy数组后问题仍未解决,请问该如何处理?

解决方法

XGBoost的survival:cox目标函数无法直接接收多列标签数据,必须通过DMatrix结构分别传递事件时间和事件状态,具体步骤如下:

  • 拆分标签字段
    先把训练集和测试集的事件时间、事件指示器分开:

    # 拆分训练集标签
    y_train_time = y_train['time_to_event']
    y_train_event = y_train['event']
    # 拆分测试集标签
    y_test_time = y_test['time_to_event']
    y_test_event = y_test['event']
    
  • 用DMatrix包装数据
    通过DMatrix的label参数传入事件时间,set_weight()方法传入事件指示器(0代表删失样本,1代表事件发生样本):

    # 构建训练集DMatrix
    dtrain = xgb.DMatrix(X_train, label=y_train_time)
    dtrain.set_weight(y_train_event)
    # 构建测试集DMatrix
    dtest = xgb.DMatrix(X_test, label=y_test_time)
    dtest.set_weight(y_test_event)
    
  • 使用原生接口训练模型
    放弃XGBRegressor的fit方法,改用xgb.train接口完成训练:

    params = {
        'objective': 'survival:cox',
        'eval_metric': 'cox-nloglik'
    }
    # 训练模型,num_boost_round可根据需求调整
    model = xgb.train(params, dtrain, num_boost_round=100, evals=[(dtest, 'test')])
    
  • 预测风险分数
    模型输出的是样本的风险分数,分数越高代表事件发生概率越高:

    y_pred = model.predict(dtest)
    

关键说明

  • XGBRegressor的封装接口不支持Cox模型所需的双标签输入,必须用原生xgb.train配合DMatrix处理。
  • 直接转Numpy数组无法解决问题,因为普通数组无法同时绑定事件时间和状态两个信息,必须通过DMatrix的专属参数传递。

内容的提问来源于stack exchange,提问作者Mohamad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 20:15:44