Scikit Pipeline搭配MultiOutputRegressor做RandomSearch参数报错求解
问题原因
- 第一次报错是参数路径和传入的搜索对象不匹配:你的完整调用链是三层嵌套结构
MultiOutputRegressor -> Pipeline -> XGBRegressor,当你把mimo_wrapper作为RandomizedSearchCV的估计器传入时,estimator__reg_alpha的路径只定位到了Pipeline层级,Pipeline本身没有reg_alpha参数,因此触发参数无效报错。 - 第二次报错是你修改参数为
estimator__estimator__reg_alpha后,误将搜索的估计器换成了pipeline_xgboost,此时Pipeline只有一层estimator(对应XGBRegressor),两层estimator的路径找不到对应参数,触发报错。 - 直接调用
set_params生效的原因是你手动先取到了mimo_wrapper.estimator(也就是Pipeline对象),此时传入estimator__max_depth刚好对应Pipeline内部的XGBRegressor参数,路径匹配因此能成功设置。
正确配置方法
选择以下任意一种方案即可:
方案1:对带多输出包装器的整体调参
保持搜索估计器为mimo_wrapper,参数路径加两层estimator前缀,分别对应MultiOutputRegressor的估计器属性、Pipeline内部的XGB组件名:
# 正确的参数定义 parameters = { 'estimator__estimator__reg_alpha': [0.0001, 0.001, 0.01, 0.1, 1, 10, 100], 'estimator__estimator__max_depth': [10, 100, 1000] # 其余XGB参数按相同前缀规则添加 } # 搜索对象传入mimo_wrapper random_grid = RandomizedSearchCV(estimator=mimo_wrapper, param_distributions=parameters, random_state=0, n_iter=5, n_jobs=-1, refit=True, cv=3, verbose=True, pre_dispatch='2*n_jobs', error_score='raise', return_train_score=True, scoring='neg_mean_absolute_error') hyperparameters_tuning = random_grid.fit(df.drop(columns=TARGETS+UMAPS), df[TARGETS])
方案2:先调Pipeline参数再套包装器
如果不需要在调参阶段关联多输出逻辑,可以先单独对Pipeline调参,参数路径只用一层estimator前缀,调参完成后再套入MultiOutputRegressor使用:
# 适配Pipeline的参数定义 parameters = { 'estimator__reg_alpha': [0.0001, 0.001, 0.01, 0.1, 1, 10, 100], 'estimator__max_depth': [10, 100, 1000] } # 搜索对象传入pipeline_xgboost random_grid = RandomizedSearchCV(estimator=pipeline_xgboost, param_distributions=parameters, random_state=0, n_iter=5, n_jobs=-1, refit=True, cv=3, verbose=True, pre_dispatch='2*n_jobs', error_score='raise', return_train_score=True, scoring='neg_mean_absolute_error') # 调参完成后再套多输出包装器 best_pipeline = random_grid.fit(df.drop(columns=TARGETS+UMAPS), df[TARGETS]).best_estimator_ mimo_wrapper = MultiOutputRegressor(best_pipeline)
内容的提问来源于stack exchange,提问作者tfkLSTM
相关产品推荐
相关产品推荐

