You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用RandomizedSearchCV出现NotFittedError:所有估计器拟合失败

问题描述

处理二分类任务时,尝试使用RandomizedSearchCV对GradientBoostingClassifier进行超参数调优,运行代码时抛出NotFittedError,提示所有估计器拟合失败。

运行代码

# Load packages
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.model_selection import GridSearchCV, RandomizedSearchCV
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.metrics import make_scorer, accuracy_score
from scipy.stats import uniform
import pandas as pd
import numpy as np
import time

import warnings
warnings.filterwarnings('ignore')
pd.set_option("display.max_columns", None)

# Make scorer: accuracy
acc_score = make_scorer(accuracy_score)

# Load dataset
trainSet = pd.read_csv('../input/train.csv')
testSet = pd.read_csv('../input/test.csv')
submitSet = pd.read_csv('../input/sample_submission.csv')

trainSet.head()

# Remove not used variables
train = trainSet.drop(columns=['Name', 'Ticket'])
train['Cabin_letter'] = train['Cabin'].str[0:1]
train['Cabin_no'] = train['Cabin'].str[1:]

train.head()

# Feature generation: training data
train = trainSet.drop(columns=['Name', 'Ticket', 'Cabin'])
train = train.dropna(axis=0)
train = pd.get_dummies(train)

train.head()

# train validation split
X_train, X_val, y_train, y_val = train_test_split(train.drop(columns=['PassengerId','Survived'], axis=0),
                                                  train['Survived'],
                                                  test_size=0.2, random_state=111,
                                                  stratify=train['Survived'])

# RandomizedSearhCV
param_rand = {'max_depth':uniform(3,10),
              'max_features':uniform(0.8,1),
              'learning_rate':uniform(0.01,1),
              'n_estimators':uniform(80,150),
              'subsample':uniform(0.8,1)}

rand = RandomizedSearchCV(estimator=GradientBoostingClassifier(), param_distributions=param_rand, scoring=acc_score, cv=5)

rand.fit(X_train.iloc[1:100,], y_train.iloc[1:100,])

错误信息

---------------------------------------------------------------------------
NotFittedError                            Traceback (most recent call last)
Input In [15], in <cell line: 10>()
      2 param_rand = {'max_depth':uniform(3,10),
      3               'max_features':uniform(0.8,1),
      4               'learning_rate':uniform(0.01,1),
      5               'n_estimators':uniform(80,150),
      6               'subsample':uniform(0.8,1)}
      8 rand = RandomizedSearchCV(estimator=GradientBoostingClassifier(), param_distributions=param_rand, scoring=acc_score, cv=5)
---&gt; 10 rand.fit(X_train.iloc[1:100,], y_train.iloc[1:100,])

File ~\anaconda3\lib\site-packages\sklearn\utils\validation.py:63, in _deprecate_positional_args.<locals>._inner_deprecate_positional_args.<locals>.inner_f(*args, **kwargs)
     61 extra_args = len(args) - len(all_args)
     62 if extra_args <= 0:
---&gt; 63     return f(*args, **kwargs)
     65 # extra_args > 0
     66 args_msg = ['{}={}'.format(name, arg)
     67             for name, arg in zip(kwonly_args[:extra_args],
     68                                  args[-extra_args:])]

File ~\anaconda3\lib\site-packages\sklearn\model_selection\_search.py:841, in BaseSearchCV.fit(self, X, y, groups, **fit_params)
    835     results = self._format_results(
    836         all_candidate_params, n_splits, all_out,
    837         all_more_results)
    839     return results
--&gt; 841 self._run_search(evaluate_candidates)
    843 # multimetric is determined here because in the case of a callable
    844 # self.scoring the return type is only known after calling
    845 first_test_score = all_out[0]['test_scores']

File ~\anaconda3\lib\site-packages\sklearn\model_selection\_search.py:1633, in RandomizedSearchCV._run_search(self, evaluate_candidates)
   1631 def _run_search(self, evaluate_candidates):
   1632     """Search n_iter candidates from param_distributions"""
--&gt; 1633     evaluate_candidates(ParameterSampler(
   1634         self.param_distributions, self.n_iter,
   1635         random_state=self.random_state))

File ~\anaconda3\lib\site-packages\sklearn\model_selection\_search.py:827, in BaseSearchCV.fit.<locals>.evaluate_candidates(candidate_params, cv, more_results)
    822 # For callable self.scoring, the return type is only know after
    823 # calling. If the return type is a dictionary, the error scores
    824 # can now be inserted with the correct key. The type checking
    825 # of out will be done in `_insert_error_scores`.
    826 if callable(self.scoring):
--&gt; 827     _insert_error_scores(out, self.error_score)
    828 all_candidate_params.extend(candidate_params)
    829 all_out.extend(out)

File ~\anaconda3\lib\site-packages\sklearn\model_selection\_validation.py:301, in _insert_error_scores(results, error_score)
    298         successful_score = result["test_scores"]
    300 if successful_score is None:
--&gt; 301     raise NotFittedError("All estimators failed to fit")
    303 if isinstance(successful_score, dict):
    304     formatted_error = {name: error_score for name in successful_score}

NotFittedError: All estimators failed to fit
错误原因
  • 超参数类型不匹配:GradientBoostingClassifier要求max_depth和n_estimators为整数,但代码中用uniform生成了浮点数,导致模型初始化失败。
  • 超参数范围无效:
    • scipy.stats.uniform的参数是loc(起始值)和scale(取值范围长度),不是上下限。比如uniform(0.8,1)会生成0.8到1.8之间的值,而max_features的有效范围是(0,1],超出范围会引发错误。
    • learning_rate的uniform(0.01,1)会生成0.01到1.01之间的值,过大的学习率会导致模型难以收敛。
  • 训练样本量过小:仅使用99条样本做5折交叉验证,每个折的样本量不足20,模型无法有效拟合,导致所有候选参数都拟合失败。
解决方法

1. 修正超参数分布

针对整数参数使用randint分布,调整所有参数到有效范围:

from scipy.stats import randint, uniform

param_rand = {
    'max_depth': randint(3, 13),  # 生成3-12的整数
    'max_features': uniform(0.6, 0.4),  # 生成0.6-1.0的浮点数
    'learning_rate': uniform(0.01, 0.2),  # 生成0.01-0.21的浮点数,避免过大学习率
    'n_estimators': randint(80, 230),  # 生成80-229的整数
    'subsample': uniform(0.7, 0.3)  # 生成0.7-1.0的浮点数
}

2. 使用完整训练样本

去掉样本切片,用全部训练数据进行拟合:

rand.fit(X_train, y_train)

3. 排查具体错误(可选)

关闭警告过滤,查看每个估计器的具体拟合错误,进一步定位问题:

warnings.filterwarnings('default')

内容的提问来源于stack exchange,提问作者UseR10085

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 02:20:50