使用RandomizedSearchCV出现NotFittedError:所有估计器拟合失败
问题描述
处理二分类任务时,尝试使用RandomizedSearchCV对GradientBoostingClassifier进行超参数调优,运行代码时抛出NotFittedError,提示所有估计器拟合失败。
运行代码
# Load packages from sklearn.model_selection import train_test_split, cross_val_score from sklearn.model_selection import GridSearchCV, RandomizedSearchCV from sklearn.ensemble import GradientBoostingClassifier from sklearn.metrics import make_scorer, accuracy_score from scipy.stats import uniform import pandas as pd import numpy as np import time import warnings warnings.filterwarnings('ignore') pd.set_option("display.max_columns", None) # Make scorer: accuracy acc_score = make_scorer(accuracy_score) # Load dataset trainSet = pd.read_csv('../input/train.csv') testSet = pd.read_csv('../input/test.csv') submitSet = pd.read_csv('../input/sample_submission.csv') trainSet.head() # Remove not used variables train = trainSet.drop(columns=['Name', 'Ticket']) train['Cabin_letter'] = train['Cabin'].str[0:1] train['Cabin_no'] = train['Cabin'].str[1:] train.head() # Feature generation: training data train = trainSet.drop(columns=['Name', 'Ticket', 'Cabin']) train = train.dropna(axis=0) train = pd.get_dummies(train) train.head() # train validation split X_train, X_val, y_train, y_val = train_test_split(train.drop(columns=['PassengerId','Survived'], axis=0), train['Survived'], test_size=0.2, random_state=111, stratify=train['Survived']) # RandomizedSearhCV param_rand = {'max_depth':uniform(3,10), 'max_features':uniform(0.8,1), 'learning_rate':uniform(0.01,1), 'n_estimators':uniform(80,150), 'subsample':uniform(0.8,1)} rand = RandomizedSearchCV(estimator=GradientBoostingClassifier(), param_distributions=param_rand, scoring=acc_score, cv=5) rand.fit(X_train.iloc[1:100,], y_train.iloc[1:100,])
错误信息
--------------------------------------------------------------------------- NotFittedError Traceback (most recent call last) Input In [15], in <cell line: 10>() 2 param_rand = {'max_depth':uniform(3,10), 3 'max_features':uniform(0.8,1), 4 'learning_rate':uniform(0.01,1), 5 'n_estimators':uniform(80,150), 6 'subsample':uniform(0.8,1)} 8 rand = RandomizedSearchCV(estimator=GradientBoostingClassifier(), param_distributions=param_rand, scoring=acc_score, cv=5) ---> 10 rand.fit(X_train.iloc[1:100,], y_train.iloc[1:100,]) File ~\anaconda3\lib\site-packages\sklearn\utils\validation.py:63, in _deprecate_positional_args.<locals>._inner_deprecate_positional_args.<locals>.inner_f(*args, **kwargs) 61 extra_args = len(args) - len(all_args) 62 if extra_args <= 0: ---> 63 return f(*args, **kwargs) 65 # extra_args > 0 66 args_msg = ['{}={}'.format(name, arg) 67 for name, arg in zip(kwonly_args[:extra_args], 68 args[-extra_args:])] File ~\anaconda3\lib\site-packages\sklearn\model_selection\_search.py:841, in BaseSearchCV.fit(self, X, y, groups, **fit_params) 835 results = self._format_results( 836 all_candidate_params, n_splits, all_out, 837 all_more_results) 839 return results --> 841 self._run_search(evaluate_candidates) 843 # multimetric is determined here because in the case of a callable 844 # self.scoring the return type is only known after calling 845 first_test_score = all_out[0]['test_scores'] File ~\anaconda3\lib\site-packages\sklearn\model_selection\_search.py:1633, in RandomizedSearchCV._run_search(self, evaluate_candidates) 1631 def _run_search(self, evaluate_candidates): 1632 """Search n_iter candidates from param_distributions""" --> 1633 evaluate_candidates(ParameterSampler( 1634 self.param_distributions, self.n_iter, 1635 random_state=self.random_state)) File ~\anaconda3\lib\site-packages\sklearn\model_selection\_search.py:827, in BaseSearchCV.fit.<locals>.evaluate_candidates(candidate_params, cv, more_results) 822 # For callable self.scoring, the return type is only know after 823 # calling. If the return type is a dictionary, the error scores 824 # can now be inserted with the correct key. The type checking 825 # of out will be done in `_insert_error_scores`. 826 if callable(self.scoring): --> 827 _insert_error_scores(out, self.error_score) 828 all_candidate_params.extend(candidate_params) 829 all_out.extend(out) File ~\anaconda3\lib\site-packages\sklearn\model_selection\_validation.py:301, in _insert_error_scores(results, error_score) 298 successful_score = result["test_scores"] 300 if successful_score is None: --> 301 raise NotFittedError("All estimators failed to fit") 303 if isinstance(successful_score, dict): 304 formatted_error = {name: error_score for name in successful_score} NotFittedError: All estimators failed to fit
错误原因
- 超参数类型不匹配:
GradientBoostingClassifier要求max_depth和n_estimators为整数,但代码中用uniform生成了浮点数,导致模型初始化失败。 - 超参数范围无效:
scipy.stats.uniform的参数是loc(起始值)和scale(取值范围长度),不是上下限。比如uniform(0.8,1)会生成0.8到1.8之间的值,而max_features的有效范围是(0,1],超出范围会引发错误。learning_rate的uniform(0.01,1)会生成0.01到1.01之间的值,过大的学习率会导致模型难以收敛。
- 训练样本量过小:仅使用99条样本做5折交叉验证,每个折的样本量不足20,模型无法有效拟合,导致所有候选参数都拟合失败。
解决方法
1. 修正超参数分布
针对整数参数使用randint分布,调整所有参数到有效范围:
from scipy.stats import randint, uniform param_rand = { 'max_depth': randint(3, 13), # 生成3-12的整数 'max_features': uniform(0.6, 0.4), # 生成0.6-1.0的浮点数 'learning_rate': uniform(0.01, 0.2), # 生成0.01-0.21的浮点数,避免过大学习率 'n_estimators': randint(80, 230), # 生成80-229的整数 'subsample': uniform(0.7, 0.3) # 生成0.7-1.0的浮点数 }
2. 使用完整训练样本
去掉样本切片,用全部训练数据进行拟合:
rand.fit(X_train, y_train)
3. 排查具体错误(可选)
关闭警告过滤,查看每个估计器的具体拟合错误,进一步定位问题:
warnings.filterwarnings('default')
内容的提问来源于stack exchange,提问作者UseR10085
相关产品推荐
相关产品推荐

