You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Dask RandomizedSearchCV调优XGBoost报sample_weight不支持错误

问题描述

使用Dask与RandomizedSearchCV对XGBoost Regressor模型进行超参数调优时,抛出异常Exception: 'ValueError("\'sample_weight\' is not supported.")',代码中未手动使用sample_weight参数,无法定位错误触发原因。

报错日志

2022-06-25 10:12:06,128 - distributed.worker - WARNING - Compute Failed
Key:       ('xgbregressor-fit-score-68a8811442e8db26da73c496969ce684', 0, 0)
Function:  fit_and_score
args:      (XGBRegressor(base_score=None, booster=None, colsample_bylevel=None,
             colsample_bynode=None, colsample_bytree=None,
             enable_categorical=False, gamma=None, gpu_id=None,
             importance_type=None, interaction_constraints=None,
             learning_rate=None, max_delta_step=None, max_depth=None,
             min_child_weight=None, missing=-999.0, monotone_constraints=None,
             n_estimators=100, n_jobs=1, num_parallel_tree=None, predictor=None,
             random_state=0, reg_alpha=None, reg_lambda=None,
             scale_pos_weight=None, subsample=None, tree_method='gpu_hist',
             validate_parameters=None, verbosity=None),
Exception: 'ValueError("\'sample_weight\' is not supported.")'

问题复现代码

# Define our model
cluster = LocalCUDACluster(dashboard_address="127.0.0.1:8005")
client = Client(cluster)
params_fixed = {'objective'   : 'reg:squarederror', 
            'random_state': 0,
            'n_jobs'      : 1,
            'tree_method' : 'gpu_hist',
            'missing' : -999.0
        }
params_hyp = {
'n_estimators': [500, 800, 1000],
'max_depth':[5,7, 10, 12],
'min_child_weight': [0.9, 1.0],
'subsample': [0.9, 1.0],
'colsample_bylevel': [0.9, 1.0],
'colsample_bynode': [0.9, 1.0],
'colsample_bytree': [0.9, 1.0]}


regressor = xgb.XGBRegressor(**params_fixed)

def do_HPO(model, gridsearch_params, scorer, X, y, mode='gpu-Grid', n_iter=10):
    """
        Perform HPO based on the mode specified
        
        mode: default gpu-Grid. The possible options are:
        1. gpu-grid: Perform GPU based GridSearchCV
        2. gpu-random: Perform GPU based RandomizedSearchCV
        
        n_iter: specified with Random option for number of parameter settings sampled
        
        Returns the best estimator and the results of the search
    """
    if mode == 'gpu-grid':
        print("gpu-grid selected")
        clf = dcv.GridSearchCV(model,
                               gridsearch_params,
                               cv=N_FOLDS,
                               scoring=scorer)
    elif mode == 'gpu-random':
        print("gpu-random selected")
        clf = dcv.RandomizedSearchCV(model,
                               gridsearch_params,
                               cv=N_FOLDS,
                               scoring=scorer,
                               n_iter=n_iter)

    else:
        print("Unknown Option, please choose one of [gpu-grid, gpu-random]")
        return None, None
    res = clf.fit(X, y)
    print("Best clf and score {} {}\n---\n".format(res.best_estimator_, res.best_score_))
    return res.best_estimator_, res
    mode = "gpu-random"
    

res, results = do_HPO(regressor,
                                params_hyp,
                                mean_absolute_error,
                                X,
                                y,
                                mode=mode,
                                n_iter=N_ITER)
错误原因
  • 直接将裸的mean_absolute_error函数传入scoring参数,未做包装。Dask-ML的交叉验证内置逻辑在识别到未通过make_scorer包装的原生指标函数时,会自动启用样本权重传递逻辑,在fit_and_score环节隐式向模型传入sample_weight参数,不需要用户手动指定。
  • 代码使用tree_method='gpu_hist'的GPU模式XGBoost,2022年中早期版本的dask-ml、dask-xgboost与XGBoost的适配层存在缺陷,GPU训练接口没有实现sample_weight参数的接收逻辑,收到隐式传入的该参数时直接抛出不支持的错误。
  • 代码中return语句后写的mode = "gpu-random"是永远不会执行的死代码,和本次报错无关。
解决方法
  • 首先将传入的评分指标用sklearn.metrics.make_scorer包装,阻断Dask CV隐式传递sample_weight的逻辑,修改示例:
    from sklearn.metrics import make_scorer, mean_absolute_error
    # 回归任务MAE是越小越好,所以greater_is_better设为False
    mae_scorer = make_scorer(mean_absolute_error, greater_is_better=False)
    # 调用do_HPO时把原来的mean_absolute_error换成mae_scorer
    res, results = do_HPO(regressor,
                                    params_hyp,
                                    mae_scorer,
                                    X,
                                    y,
                                    mode=mode,
                                    n_iter=N_ITER)
    
  • 如果上述修改后仍报错,可以在初始化GridSearchCV/RandomizedSearchCV时显式添加参数return_train_score=False,并在调用fit时传入fit_params={},显式声明不传入任何额外拟合参数,彻底阻断sample_weight的隐式传递路径。
  • 若仍存在兼容问题,可将xgboost、dask-ml、dask-xgboost三个库升级到2022年6月之后的稳定版本,后续版本已经修复了GPU模式下sample_weight参数的适配问题。
  • 可删除do_HPO函数中return语句后的无效死代码,避免逻辑误导。

内容的提问来源于stack exchange,提问作者John Doe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 06:27:14