You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XGBoost二分类任务超参数调参报错求助

糖尿病二分类任务XGBoost贝叶斯调参报错解决

问题背景

使用XGBoost和scikit-learn进行糖尿病患者二分类(患病/未患病),标签列"Outcome"取值为0或1。常规训练流程正常,但使用BayesSearchCV做超参数调参时触发报错。

错误信息

File "D:\code\python\ml\venv\lib\site-packages\xgboost\core.py", line 1803, in eval_set
  _check_call(
File "D:\code\python\ml\venv\lib\site-packages\xgboost\core.py", line 203, in _check_call
  raise XGBoostError(py_str(_LIB.XGBGetLastError()))
xgboost.core.XGBoostError: [19:20:51] C:/Users/Administrator/workspace/xgboost-win64_release_1.6.0/src/metric/multiclass_metric.cu:35: Check failed: label_error >= 0 && label_error < static_cast<int32_t>(n_class): MultiClassEvaluation: label must be in [0, num_class), num_class=1 but found 1 in label

原代码

import pandas as pd
import numpy as np
import xgboost as xgb
from sklearn.model_selection import train_test_split
from skopt import BayesSearchCV
from sklearn.model_selection import KFold

diabetes2_path = r"D:\code\python\ml\datasets\classification\diabetes(2).csv"

df = pd.read_csv(diabetes2_path)
X_train, X_test, y_train, y_test = train_test_split(df, df["Outcome"], test_size=0.25)

params = {'max_depth': [3, 4, 5, 6, 7, 8, 9, 10, 15, 20],
          'learning_rate': [0.01, 0.05, 0.1, 0.15, 0.2],
          'subsample': np.arange(0.5, 1.0, 0.1),
          'colsample_bytree': np.arange(0.4, 1.0, 0.1),
          'colsample_bylevel': np.arange(0.4, 1.0, 0.1),
          'min_child_weight': (0, 30),
          'gamma': (0, 20),
          'num_class': np.arange(2),
          'n_estimators': (50, 1000)}

xgbr = xgb.XGBClassifier(enable_categorical=True, tree_method='hist')
# xgbr.fit(X_train, y_train)
# result = xgbr.predict(X_test) # works great!

kfold = KFold(n_splits=3)
search = BayesSearchCV(estimator=xgbr,
                       search_spaces=params,
                       scoring='neg_mean_squared_error',
                       n_iter=25,
                       verbose=0)

eval_set = [(X_test, y_test)]
result = search.fit(X_train, y_train, early_stopping_rounds=10, eval_metric="mlogloss",
                    eval_set=eval_set)  # ERROR!!

错误原因及修正方案

1. 错误设置num_class参数

二分类任务无需手动指定num_class,XGBoost会自动根据标签取值识别分类数。原参数中num_class: np.arange(2)会生成0和1两个候选值,当调参抽到num_class=1时,XGBoost判定为单分类任务,但标签中存在1,超出[0,1)的合法范围,触发报错。
修正:从params字典中删除'num_class': np.arange(2)。

2. 训练数据拆分错误

train_test_split(df, df["Outcome"])将包含标签列的整个数据集作为特征输入,导致特征与标签重复,干扰模型训练逻辑。
修正:拆分时用df.drop("Outcome", axis=1)提取纯特征集:

X = df.drop("Outcome", axis=1)
y = df["Outcome"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25)

3. 评估指标不匹配

eval_metric="mlogloss"是多分类任务的专属评估指标,二分类任务应使用"logloss"(对应二分类对数损失)。
修正:将eval_metric改为"logloss"。

4. 早停参数传递方式错误

BayesSearchCV的fit方法中,early_stopping_rounds、eval_set等参数需要通过fit_params字典统一传递,否则会导致参数传递异常。
修正:将早停相关参数放入fit_params后传递:

fit_params = {
    "early_stopping_rounds": 10,
    "eval_metric": "logloss",
    "eval_set": [(X_test, y_test)],
    "verbose": 0
}
result = search.fit(X_train, y_train, **fit_params)

修正后的完整代码

import pandas as pd
import numpy as np
import xgboost as xgb
from sklearn.model_selection import train_test_split
from skopt import BayesSearchCV
from sklearn.model_selection import KFold

diabetes2_path = r"D:\code\python\ml\datasets\classification\diabetes(2).csv"

df = pd.read_csv(diabetes2_path)
# 正确拆分特征与标签
X = df.drop("Outcome", axis=1)
y = df["Outcome"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25)

# 删除错误的num_class参数
params = {'max_depth': [3, 4, 5, 6, 7, 8, 9, 10, 15, 20],
          'learning_rate': [0.01, 0.05, 0.1, 0.15, 0.2],
          'subsample': np.arange(0.5, 1.0, 0.1),
          'colsample_bytree': np.arange(0.4, 1.0, 0.1),
          'colsample_bylevel': np.arange(0.4, 1.0, 0.1),
          'min_child_weight': (0, 30),
          'gamma': (0, 20),
          'n_estimators': (50, 1000)}

xgbr = xgb.XGBClassifier(enable_categorical=True, tree_method='hist')

kfold = KFold(n_splits=3)
search = BayesSearchCV(estimator=xgbr,
                       search_spaces=params,
                       scoring='neg_mean_squared_error',
                       n_iter=25,
                       verbose=0)

# 用fit_params传递早停相关参数
fit_params = {
    "early_stopping_rounds": 10,
    "eval_metric": "logloss",
    "eval_set": [(X_test, y_test)],
    "verbose": 0
}
result = search.fit(X_train, y_train, **fit_params)

内容的提问来源于stack exchange,提问作者Zag Gol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 17:06:20