使用Hyperopt调参RandomForestClassifier时max_features参数报错求助
问题:scikit-learn最新版RandomForestClassifier与Hyperopt超参调优报错
问题现象
使用最新版scikit-learn和Hyperopt进行随机森林模型的超参数调优时,设置max_features参数可选值为['auto','sqrt','log2',None],触发InvalidParameterError,报错信息如下:
InvalidParameterError: The 'max_features' parameter of RandomForestClassifier must be an int in the range [1, inf), a float in the range (0.0, 1.0], a str among {'sqrt', 'log2'} or None. Got 'auto' instead.
注释max_features参数后代码可正常运行。
报错原因
scikit-learn在较新版本中已移除RandomForestClassifier的max_features参数的'auto'选项,当前合法取值范围为:
- 整数:≥1的整数,表示每次分裂时考虑的特征数量
- 浮点数:(0.0,1.0]区间内的浮点数,表示每次分裂时考虑的特征占总特征数的比例
- 字符串:仅接受
'sqrt'或'log2' - None:等价于旧版本的
'auto',即每次分裂时考虑所有特征
修复方案
将Hyperopt搜索空间中max_features的可选值里的'auto'移除,保留['sqrt','log2',None]即可,None完全替代原'auto'的功能。
修改后的完整代码
from hyperopt import hp, fmin, tpe, STATUS_OK, Trials from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import cross_val_score # 假设X_train、Y_train已提前定义 space={ 'criterion': hp.choice('criterion',['entropy','gini']), 'max_depth': hp.quniform('max_depth',10,1200,10), 'max_features': hp.choice('max_features',['sqrt','log2',None]), # 移除'auto'选项 'min_samples_leaf': hp.uniform('min_samples_leaf',0,0.5), 'min_samples_split': hp.uniform('min_samples_split',0,1), 'n_estimators': hp.choice('n_estimators',[10,50,300,750,1200,1300,1800,2000]) } def objective(space): model=RandomForestClassifier( criterion=space['criterion'], max_depth=int(space['max_depth']), max_features=space['max_features'], min_samples_leaf=space['min_samples_leaf'], min_samples_split=space['min_samples_split'], n_estimators=space['n_estimators'] ) accuracy=cross_val_score(model,X_train,Y_train,cv=5).mean() return {'loss':-accuracy,'status':STATUS_OK} trials=Trials() best=fmin( fn=objective, space=space, algo=tpe.suggest, max_evals=80, trials=trials ) print(best)
内容的提问来源于stack exchange,提问作者Abhinav Khandelwal
相关产品推荐
相关产品推荐

