You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

心脏病分类任务中GridSearchCV调优RandomForestClassifier遇拟合失败

心脏病分类任务中RandomForest超参数调优失败问题排查

问题描述

我正在处理一项心脏病相关的分类任务,使用RandomForestClassifier模型。在对该模型进行超参数调优时遇到问题,我使用sklearn的Pipeline和ColumnTransformer进行预处理。

错误信息

Error: 720 fits failed out of a total of 2160.
The score on these train-test partitions for these parameters will be set to nan.
If these failures are not expected, you can try to debug them by setting error_score='raise'.
UserWarning: One or more of the test scores are non-finite

相关代码

numerical_pipeline = Pipeline(
steps=[('scaler',StandardScaler())]
)

categorical_pipeline = Pipeline(
steps=[('encoder',OneHotEncoder(handle_unknown='ignore'))]  
)

preprocessor = ColumnTransformer(
[('numerical_pipeline',numerical_pipeline,numerical_features),
 ('categorical_pipeline',categorical_pipeline,categorical_features)]

X_train,X_test,y_train,y_test = train_test_split(X,y,test_size=0.3)

scaled_X_train = preprocessor.fit_transform(X_train)
scaled_X_test = preprocessor.transform(X_test)

param_grid={'max_depth':[3,5,10,None],
          'n_estimators':[10,100,200],
          'max_features':[1,3,5,7],
          'min_samples_leaf':[1,2,3],
          'min_samples_split':[1,2,3]
       }

grid = GridSearchCV(RandomForestClassifier(),param_grid=param_grid,cv=5,scoring='accuracy',verbose=True,n_jobs=-1)
grid.fit(scaled_X_train,y_train)

问题原因与解决方法

  • 核心问题1:min_samples_split参数取值错误
    RandomForest的min_samples_split要求必须≥2,你设置了[1,2,3],当取1时会直接报错——分裂节点至少需要2个样本才能继续划分,这是导致大量拟合失败的主要原因。

  • 问题2:预处理与模型未整合到同一Pipeline
    先对整个训练集做预处理再传入GridSearch,会导致交叉验证时出现数据泄露(预处理的scaler/encoder用了全部训练集拟合,而非每个fold的训练子集)。正确做法是把预处理和模型整合进同一个Pipeline,让GridSearch在每个fold内部完成拟合-变换流程。

修正后的代码

from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split, GridSearchCV

# 构建完整Pipeline,整合预处理与模型
full_pipeline = Pipeline([
    ('preprocessor', ColumnTransformer([
        ('numerical', StandardScaler(), numerical_features),
        ('categorical', OneHotEncoder(handle_unknown='ignore'), categorical_features)
    ])),
    ('classifier', RandomForestClassifier())
])

# 修正参数网格,移除min_samples_split=1
param_grid={
    'classifier__max_depth':[3,5,10,None],
    'classifier__n_estimators':[10,100,200],
    'classifier__max_features':[1,3,5,7],
    'classifier__min_samples_leaf':[1,2,3],
    'classifier__min_samples_split':[2,3]  # 取值≥2
}

# 拆分数据集
X_train,X_test,y_train,y_test = train_test_split(X,y,test_size=0.3)

# 运行GridSearch
grid = GridSearchCV(full_pipeline, param_grid=param_grid, cv=5, scoring='accuracy', verbose=True, n_jobs=-1)
grid.fit(X_train, y_train)

额外调试建议

如果仍有报错,给GridSearchCV加上error_score='raise',会直接抛出具体错误信息,方便定位剩余问题。

内容的提问来源于stack exchange,提问作者vijai vikram

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 19:06:32