You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GridSearchCV是否有参数可选训练与测试集差异最小的最优模型?

模型调优需求与问题

我的目标是得到拟合良好的模型(训练集与测试集指标差异控制在1%-5%),因为Random Forest极易过拟合——默认参数下类别1的训练集f1得分达到了1.0。

当前核心问题:GridSearchCV仅关注测试集指标,完全忽略训练集指标,导致最终选出的模型仍处于过拟合状态。

已尝试的操作

  • 尝试访问cv_results_属性,但输出内容过于繁杂,无法解读,且认为不应手动处理这些数据。

相关代码

# model definition
rf_cv = GridSearchCV(estimator=rf_clf_default,
                     # 关注模型识别类别1的能力
                     scoring=make_scorer(score_func=f1_score, pos_label=1),
                     param_grid={'randomforestclassifier__n_estimators': [37,38,39,100,200],
                                 'randomforestclassifier__max_depth': [4,5,6,10,20,30],
                                 'randomforestclassifier__min_samples_leaf': [2,3,4]},
                     return_train_score=True,
                     refit=True)

# 忽略OneHotEncoder关于未知类别的警告
with warnings.catch_warnings():
    warnings.simplefilter(action="ignore", category=UserWarning)
    # 训练模型
    rf_cv.fit(X=X_train, y=y_train)

# 获取类别1的最优召回率
print("best recall score class 1", rf_cv.best_score_)

# 获取最优参数
display("best parameters", rf_cv.best_params_)

GridSearchCV选出的最优参数

GridSearchCV选出的最优参数


内容的提问来源于stack exchange,提问作者Jason Rich Darmawan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 15:05:30