You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Hyperopt对Scikit-learn VotingClassifier的子模型超参数进行调优?

用Hyperopt调优Scikit-learn VotingClassifier的子模型超参数

我明白你现在卡在哪了——单个模型用Hyperopt调优顺手得很,但碰到VotingClassifier这种带多个子模型的集成器,就不知道怎么把那些带前缀的超参数分配给对应的子模型了。别慌,核心思路就是把Hyperopt生成的带前缀参数拆出来,分别更新每个子模型的配置,再重新构建VotingClassifier,下面是具体的实现步骤和代码:

1. 修正超参空间定义

先给你的超参字典做两处调整:

  • 删掉AB__base_estimator:这是个模型对象,不属于超参数范畴;如果要调base_estimator的参数,得用嵌套前缀(比如AB__base_estimator__max_depth),不过按你当前需求,我们先固定base_estimator的基础结构,只调AdaBoost自身的参数。
  • 统一超参数命名:比如AdaBoost的n_estimators要加上AB__前缀,避免和其他模型的参数混淆。

修正后的超参空间:

from hyperopt import hp

dic_clf_params = {
    # 决策树子模型参数,前缀DT__对应后续estimators里的命名
    'DT__max_depth': hp.choice('DT__max_depth', range(2, 20)),
    'DT__criterion': hp.choice('DT__criterion', ['gini', 'entropy']),
    'DT__class_weight': hp.choice('DT__class_weight', ['balanced']),
    # AdaBoost子模型参数,前缀AB__对应后续estimators里的命名
    'AB__n_estimators': hp.choice('AB__n_estimators', range(2, 1000)),
    'AB__learning_rate': hp.uniform('AB__learning_rate', 0.2, 1)
}

2. 重写目标函数

目标函数的核心是拆分带前缀的参数,把它们分配给对应的子模型,再构建VotingClassifier:

  • 先初始化基础子模型列表(注意每个子模型的名字要和超参前缀完全一致,比如'DT'、'AB')
  • 遍历Hyperopt传入的params,按__拆分前缀和参数名,更新对应子模型的配置
  • 用更新后的子模型构建软投票分类器,再做交叉验证计算F1值

完整的目标函数代码:

from sklearn.ensemble import VotingClassifier, AdaBoostClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import cross_val_score
from hyperopt import STATUS_OK

def objective_function(params):
    # 1. 初始化基础子模型(名字要和超参前缀严格对应)
    base_estimators = [
        ('DT', DecisionTreeClassifier()),
        ('AB', AdaBoostClassifier(base_estimator=DecisionTreeClassifier(max_depth=20)))
    ]
    
    # 2. 拆分超参数,更新对应子模型的配置
    updated_estimators = []
    for name, clf in base_estimators:
        # 提取当前子模型的所有超参数(比如前缀为DT__的参数)
        clf_params = {k.split('__')[1]: v for k, v in params.items() if k.startswith(f"{name}__")}
        # 更新子模型的参数
        clf.set_params(**clf_params)
        updated_estimators.append((name, clf))
    
    # 3. 构建软投票分类器
    voting_clf = VotingClassifier(estimators=updated_estimators, voting='soft')
    
    # 4. 交叉验证计算F1值,返回负F1作为loss(Hyperopt默认最小化loss)
    f1_score = cross_val_score(voting_clf, x_data, y_data.values.tolist(), scoring='f1', cv=10).mean()
    return {"loss": -f1_score, "status": STATUS_OK}

3. 启动Hyperopt调优

最后一步和你调单个模型的流程完全一致,用fmin函数启动调优即可:

from hyperopt import fmin, tpe, Trials

trials = Trials()
best_params = fmin(
    fn=objective_function,
    space=dic_clf_params,
    algo=tpe.suggest,
    max_evals=100,
    trials=trials
)

print("最优超参数组合:", best_params)

额外说明

  • 子模型的名字(比如'DT')必须和超参前缀完全匹配,不然参数会分配错误
  • 如果之后要调base_estimator的参数(比如AdaBoost里的决策树深度),可以在超参空间里加AB__base_estimator__max_depth这种嵌套前缀,set_params会自动处理这种嵌套参数
  • 记得把VotingClassifier的voting参数设为'soft',符合你要做的软投票分类器需求

内容的提问来源于stack exchange,提问作者Md Sabbir Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 16:09:08