You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scikit-learn Stacking Classifier集成模型效果不佳问题排查

Stacking集成模型效果不如单个模型的问题排查与解决

从你的描述和代码来看,集成模型效果不及单个Logistic Regression,大概率是几个关键操作细节出了问题,下面逐一分析并给出修复方案:

1. 重复调用fit()导致模型参数被覆盖

看你的Logistic Regression调参代码:

log_regressor = GridSearchCV(Log, parameters,cv =5)
log_regressor.fit(X_train, y_train)
log_regressor.best_params_
log_regressor.fit(X_train,y_train)  # 这里是冗余且错误的操作!

第一次fit()已经完成网格搜索并得到了最优参数模型(存在log_regressor.best_estimator_中),但第二次调用fit()会让GridSearchCV重新执行完整的网格搜索,不仅浪费资源,还可能导致后续调用的模型参数混乱。更严重的是,你把log_regressor(GridSearchCV实例)作为Stacking的final_estimator,会让Stacking在训练元模型时再次对Logistic Regression做交叉验证调参,引入不必要的复杂度。

修复方案:

  • 去掉重复的fit()调用,直接使用GridSearchCV的best_estimator_作为最终模型:
log_regressor = GridSearchCV(Log, parameters,cv =5)
log_regressor.fit(X_train, y_train)
# 直接取最优模型
best_log = log_regressor.best_estimator_
accuracy89 = best_log.score(X_test,y_test)
print('Logistic Regression Accuracy -->',((accuracy89)*100))
  • Stacking的final_estimator传入best_log(最优模型实例),而非GridSearchCV对象:
stackingCLF = StackingClassifier(
    estimators = estimators, 
    verbose = 2,
    final_estimator = best_log, 
    cv=5
)

2. 基模型传入GridSearchCV实例而非最优模型

你把model3_grid、svm_regressor这些GridSearchCV对象直接放进estimators列表中,虽然GridSearchCV的predict()会默认调用best_estimator_.predict(),但这种方式可能导致Stacking在生成元特征时重复执行网格搜索,增加计算成本甚至引入误差。

修复方案:
所有基模型都使用GridSearchCV的best_estimator_:

# 以SVC为例
svm_regressor = GridSearchCV(SVC(), svc_params, cv=5)
svm_regressor.fit(X_train, y_train)
best_svc = svm_regressor.best_estimator_

# 同理处理其他模型
best_knn = model3_grid.best_estimator_
best_nb = nb_regressor.best_estimator_
best_rf = rf_classifier.best_estimator_

# 构建estimators列表
estimators = [
    ('knn', best_knn),
    ('svc', best_svc),
    ('nb', best_nb),
    ('rf', best_rf),
]

3. 数据泄露:基模型调参使用了完整训练集

你在整个X_train上做GridSearchCV调参,再将调优后的模型放入Stacking中,但Stacking本身会用cv=5交叉验证生成元特征——这意味着基模型调参时已经见过了Stacking交叉验证中的所有数据(包括验证折),导致基模型在Stacking中的泛化能力下降,最终集成效果不如单独测试时的表现。

修复方案:
使用嵌套交叉验证(Nested CV)避免数据泄露:外层交叉验证评估集成模型整体性能,内层交叉验证负责基模型和元模型的调参。示例框架:

from sklearn.model_selection import StratifiedKFold

# 外层交叉验证
outer_cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
# 内层交叉验证用于调参
inner_cv = StratifiedKFold(n_splits=3, shuffle=True, random_state=42)

# 定义基模型的调参器
def get_tuned_model(model, params):
    return GridSearchCV(model, params, cv=inner_cv, n_jobs=-1)

# 构建estimators(建议移除表现极差的NB)
estimators = [
    ('knn', get_tuned_model(KNeighborsClassifier(), knn_params)),
    ('svc', get_tuned_model(SVC(), svc_params)),
    ('rf', get_tuned_model(RandomForestClassifier(), rf_params)),
]

# 元模型也用内层CV调参
final_estimator = get_tuned_model(LogisticRegression(), log_params)

# 嵌套CV评估集成模型
stackingCLF = StackingClassifier(
    estimators=estimators,
    final_estimator=final_estimator,
    cv=inner_cv,
    passthrough=True  # 传入原始特征增强元模型
)

# 外层CV评估
scores = cross_val_score(stackingCLF, X_train, y_train, cv=outer_cv, scoring='accuracy')
print(f"嵌套CV平均准确率: {scores.mean():.4f}")

4. 弱模型拖后腿:Naive Bayes表现极差

你的单个模型结果中NB的准确率只有50%,属于严重不合格的弱模型,加入Stacking会引入噪声,拉低整体性能。

修复方案:
暂时移除NB模型观察效果;如果要保留,先排查NB调参是否合理(比如var_smoothing的搜索范围),或换用其他类型的朴素贝叶斯模型(如MultinomialNB,若数据适合)。

5. 未传入原始特征给元模型

默认StackingClassifier(passthrough=False),元模型只能用到基模型的预测结果作为特征,而你的单个Logistic Regression是直接用原始特征训练的,信息密度更高。

修复方案:
设置passthrough=True,让元模型同时使用原始特征和基模型的预测结果:

stackingCLF = StackingClassifier(
    estimators = estimators, 
    verbose = 2,
    final_estimator = best_log, 
    cv=5,
    passthrough=True  # 关键参数
)

6. 交叉验证未考虑数据分层

心脏病数据集通常类别不平衡,普通KFold可能导致验证集样本分布不均,影响调参和集成效果。

修复方案:
所有交叉验证都使用StratifiedKFold,保证每折的类别分布与整体一致:

from sklearn.model_selection import StratifiedKFold

# 调参时用StratifiedKFold
log_regressor = GridSearchCV(
    Log, 
    parameters,
    cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
)

# Stacking的cv也用StratifiedKFold
stackingCLF = StackingClassifier(
    estimators = estimators, 
    verbose = 2,
    final_estimator = best_log, 
    cv=StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
)

先按照上面的步骤逐一修复,应该能显著提升Stacking集成模型的性能。

内容的提问来源于stack exchange,提问作者Gautam Goyal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 22:58:13