如何对Bagging集成的决策树模型执行GridSearchCV同步超参数调优
超参数调优实现方案
核心误区纠正
- 你完全不需要使用Pipeline:Pipeline的作用是串联「数据预处理步骤+模型」,仅当前序步骤需要调用
fit_transform时才需要。BaggingClassifier本身就支持直接传入基分类器并单独调优两类参数,和Pipeline没有关系,你之前的思路走偏了。 - 你看到的参考方案本身就支持同时调优两类参数:param_grid里,基分类器的参数加
base_estimator__前缀,Bagging集成器本身的参数直接写参数名即可,不需要加任何前缀。
完整实现代码
from sklearn.ensemble import BaggingClassifier from sklearn.tree import DecisionTreeClassifier from sklearn.model_selection import GridSearchCV # 1. 定义参数网格,同时包含两类参数 param_grid = { # 基分类器(DecisionTree)参数,加base_estimator__前缀 'base_estimator__class_weight': ['balanced'], 'base_estimator__criterion': ['gini', 'entropy'], 'base_estimator__max_depth': [1,2,3,4,5,6], 'base_estimator__max_features': [1,'auto'], 'base_estimator__min_weight_fraction_leaf': [0.001, 0.005, 0.01], 'base_estimator__random_state': [0], 'base_estimator__splitter': ['best','random'], # Bagging集成器参数,直接写参数名 'bootstrap': [True,False], 'max_features': [1.0,2.0,3.0], 'max_samples': [avgUniqueness, 1.0], 'n_estimators': [10,50,100,1000] } # 2. 初始化Bagging分类器,固定不需要调优的参数 bagging_clf = BaggingClassifier( base_estimator=DecisionTreeClassifier(), bootstrap_features=False, # 固定值不需要加入参数网格 random_state=0 ) # 3. 初始化网格搜索 grid_search = GridSearchCV( estimator=bagging_clf, param_grid=param_grid, scoring='替换为你需要的评估指标', # 如f1、roc_auc等 cv=5, # 交叉验证折数可自行调整 n_jobs=-1 # 调用全部CPU核心加速搜索 ) # 4. 训练时传入sample_weight即可,会自动传递给基分类器 grid_search.fit(X, y, sample_weight=你的样本权重变量) # 输出最优参数、获取最优模型 print(grid_search.best_params_) best_model = grid_search.best_estimator_
注意事项
- 请提前确认你的数据集特征总数大于等于3,否则
max_features设为2.0/3.0会触发参数合法性报错,按需调整该参数的取值范围即可。 - 如果搜索速度过慢,可以适当缩小参数范围,或者改用
RandomizedSearchCV做随机搜索,在精度损失不大的前提下大幅提升搜索效率。
内容的提问来源于stack exchange,提问作者Javier C Salaverri
相关产品推荐
相关产品推荐

