在集成Pipeline中运行GridSearch后如何提取最佳参数?
提取Pipeline中GridSearchCV的最佳参数
嘿,这个问题其实很简单,我来给你一步步拆解怎么拿到你要的最佳参数:
首先,假设你已经用你的Pipeline拟合了数据(也就是运行了 pipeline_rf.fit(X, y)),接下来只需要定位到Pipeline里的GridSearchCV组件,然后调用它的内置属性就行:
步骤1:获取Pipeline中的GridSearchCV对象
因为你的GridSearchCV是Pipeline里名为grid_search_lr的步骤,所以可以通过named_steps属性来获取这个对象:
# 提取Pipeline中的GridSearchCV实例 grid_search_instance = pipeline_rf.named_steps['grid_search_lr']
步骤2:提取最佳参数
拿到GridSearchCV实例后,直接访问best_params_属性就能得到最优的参数组合,它会以字典的形式返回:
# 获取最佳参数字典 best_parameters = grid_search_instance.best_params_ print(best_parameters)
输出结果会是类似这样的结构:
{'bootstrap': True, 'max_depth': 100, 'max_features': 'sqrt', 'min_samples_leaf': 2, 'min_samples_split': 5, 'n_estimators': 1000}
额外小技巧:获取训练好的最佳模型
如果你需要直接使用已经用最佳参数训练好的RandomForest模型,可以访问best_estimator_属性,它会返回一个已经在全量训练数据上拟合好的模型实例:
# 获取训练完成的最佳随机森林模型 best_rf_model = grid_search_instance.best_estimator_ # 可以直接用这个模型做预测 predictions = best_rf_model.predict(X_test)
要注意的是,你在GridSearchCV里设置了refit=True(这也是默认值),所以best_estimator_已经是用整个训练集重新拟合好的最优模型,不需要再额外训练啦。
内容的提问来源于stack exchange,提问作者Rahul Dev
相关产品推荐
相关产品推荐

