You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用AdaBoostClassifier时预测结果仅单一类别的问题排查

AdaBoostClassifier预测仅输出单一类别问题排查

问题现象

作为Python机器学习新手,使用AdaBoostClassifier训练模型时遇到以下问题:

  • 模型预测结果仅输出单一类别(3 - 6 m)
  • 测试集准确率固定在0.416左右
  • 已尝试调整learning_rate、将标签y转换为整数、检查数据集,且在iris示例数据集上模型运行正常

相关代码

### my dataset
import pandas as pd
csv_url = 'https://raw.githubusercontent.com/ga59wig/419B/main/data.csv?token=GHSAT0AAAAAACKSAONPXCVHO2L4IGQCID72ZK3422Q'
gdf = pd.read_csv(csv_url)

gdf = gdf.dropna()

#for accuracy_score
from sklearn import metrics as metrics

# for train_test_split
from sklearn import model_selection as model_selection

# For the classifier
from sklearn import ensemble as ensemble 

cols = gdf.columns
cols = cols[1:]

x = gdf[cols].values
y = gdf["Relative Height bin98 (cm)"]

print(y)
print(x)

x_train, x_test, y_train, y_test = model_selection.train_test_split(x,y, test_size=0.2, random_state=1569)

adaBoost_model = ensemble.AdaBoostClassifier(n_estimators=200, learning_rate=1e-05)
adaBoost_model.fit(x_train, y_train)
adaBoost_prediction = adaBoost_model.predict(x_test)

adaBoost_accuracy = metrics.accuracy_score(adaBoost_prediction, y_test)
adaBoost_confusion_matrix = metrics.confusion_matrix(adaBoost_prediction, y_test)
adaBoost_classification_report = metrics.classification_report(adaBoost_prediction, y_test)

print("Accuracy:", adaBoost_accuracy)
print(adaBoost_confusion_matrix)
print(adaBoost_classification_report)

模型输出

Accuracy: 0.4168228190714137
[[8008 3144 4233 3827]
 [   0    0    0    0]
 [   0    0    0    0]
 [   0    0    0    0]]
              precision    recall  f1-score   support

     3 - 6 m       1.00      0.42      0.59     19212
    6 - 10 m       0.00      0.00      0.00         0
        <3 m       0.00      0.00      0.00         0
      > 10 m       0.00      0.00      0.00         0

    accuracy                           0.42     19212
   macro avg       0.25      0.10      0.15     19212
weighted avg       1.00      0.42      0.59     19212

问题排查与解决方法

1. 检查数据集类别分布

首先确认训练集的类别是否严重不平衡:

  • 执行print(y_train.value_counts())查看各类别样本数量,如果3 - 6 m占比极高(接近当前准确率42%),模型会默认倾向于预测多数类来保证基础准确率。

2. 调整AdaBoost参数与基础分类器

AdaBoost默认使用深度为1的决策树(决策桩),表达能力有限,面对不平衡数据更容易偏向多数类:

  • 更换基础分类器,例如设置base_estimator=DecisionTreeClassifier(max_depth=3),提升模型复杂度
  • 调整learning_rate:当前1e-05过小,模型更新缓慢,建议尝试0.1到1之间的值
  • 增加n_estimators:若降低学习率,可同步增加弱分类器数量(如设置为500)

3. 处理类别不平衡

若确认存在类别不平衡,可通过以下方式优化:

  • 设置class_weight='balanced',让模型自动给少数类分配更高权重
  • 对少数类进行过采样(如SMOTE),或对多数类进行欠采样
  • 改用F1-score、AUC-ROC等更适合不平衡数据的评估指标,替代单一的准确率

4. 特征优化

  • 执行print(adaBoost_model.feature_importances_)查看特征重要性,移除无贡献的冗余特征
  • 若更换非树模型作为基础分类器,需先对特征进行标准化处理

内容的提问来源于stack exchange,提问作者Jonas Apelt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 15:53:15