You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义AdaBoost分类器使用LearningCurveDisplay绘制学习曲线报错求助

自定义AdaBoost分类器绘制学习曲线的解决方案

问题背景

为完成学校项目,我开发了一个以决策树桩(DecisionStump)作为弱学习器的AdaBoost分类器,具体实现代码如下:

class AdaBoost:
    def __init__(self, boosting_rounds):
        self.boosting_rounds = boosting_rounds  # number of weak learners used
        self.weak_learners = []  # [ (learner, amount of say), ... ]
        self.decision_stumps = []

    def fit(self, x, y):
        n_rows = x.shape[0]  # number of examples in training set
        attributes = x.columns  # features in training set

        # dataset example weights
        d = np.full(n_rows, 1 / n_rows)

        # preparing all decision stumps
        for a in attributes:
            # finding possible thresholds for decision stump
            values = x[a].unique()
            values.sort()
            thresholds = []
            for i in range(1, len(values)):
                thresholds.append((values[i] + values[i - 1]) / 2)
            for threshold in thresholds:
                self.decision_stumps.append(DecisionStump(a, threshold))

        for t in range(0, self.boosting_rounds):

            # choosing decision stump that minimizes weights dependent error
            min_error = float("inf")
            h = None

            for ds in self.decision_stumps:
                ds_error = 0
                for example in range(0, n_rows):
                    example_row = x.iloc[example]
                    if ds.predict(example_row) != y.iloc[example][0]:  # misclassification
                        ds_error += d[example]
                if ds_error < min_error:
                    min_error = ds_error
                    h = ds

            # amount of say
            small = 1e-10
            alpha = 0.5 * np.log((1 - min_error) / (min_error + small))  # small avoids division by 0

            # storing learner with corresponding amount of say
            self.weak_learners.append((h, alpha))
            
            # updating weights
            for i in range(0, len(d)):
                d[i] = d[i] * np.exp(-alpha * y.iloc[i] * h.predict(x.iloc[i]))
            z = np.sum(d)  # normalisation factor
            for i in range(0, len(d)):
                d[i] /= z

    def predict(self, x):
        s = 0  # sum of predictions
        for i in range(0, len(self.weak_learners)):
            prediction = self.weak_learners[i][0].predict(x)
            s += prediction * self.weak_learners[i][1]  # prediction * amount of say

        return int(np.sign(s))

尝试使用sklearn.model_selection.LearningCurveDisplay绘制学习曲线时,先是报错提示估算器没有评分函数,添加评分函数后又出现其他问题,需要找到正确的绘制方法。

解决步骤

1. 让自定义AdaBoost符合sklearn估算器规范

sklearn工具类要求估算器要么实现score方法,要么在调用时指定scoring参数。给AdaBoost类添加score方法,计算分类准确率:

from sklearn.metrics import accuracy_score
import numpy as np
import pandas as pd

class AdaBoost:
    # 保留原有__init__、fit、predict方法...
    
    def score(self, X, y):
        # 处理输入:如果y是DataFrame,转成一维数组
        if isinstance(y, pd.DataFrame):
            y = y.iloc[:, 0].values
        # 批量预测
        predictions = [self.predict(X.iloc[i]) for i in range(X.shape[0])]
        return accuracy_score(y, predictions)

2. 调整训练数据的y格式

你的fit方法中使用y.iloc[i][0],说明y是DataFrame格式,训练前将y转换为一维数组或Series,避免后续出错:

# 假设原始数据是X_train(DataFrame)、y_train(DataFrame)
y_train = y_train.iloc[:, 0].values  # 转成一维numpy数组

3. 正确绘制学习曲线

使用LearningCurveDisplay.from_estimator时,确保参数符合要求,示例代码如下:

from sklearn.model_selection import LearningCurveDisplay
import matplotlib.pyplot as plt

# 初始化自定义AdaBoost分类器
adaboost_clf = AdaBoost(boosting_rounds=10)

# 绘制学习曲线
display = LearningCurveDisplay.from_estimator(
    estimator=adaboost_clf,
    X=X_train,
    y=y_train,
    train_sizes=np.linspace(0.1, 1.0, 5),  # 训练集比例
    cv=5,  # 5折交叉验证
    scoring="accuracy",  # 显式指定评分指标(可选,已实现score方法时可省略)
    n_jobs=-1,  # 使用所有CPU核心加速
    line_kw={"marker": "o"},
    random_state=42
)

plt.title("Learning Curve for Custom AdaBoost Classifier")
plt.show()

额外问题排查

  • 决策树桩兼容性:确保DecisionStump类的predict方法能正确处理单条样本(如DataFrame的行),返回值为1或-1(与AdaBoost标签格式匹配)。
  • 标签格式校验:训练数据的y标签必须是1和-1,而非0和1,否则会影响权重更新逻辑。
  • 性能优化:当前fit方法遍历所有决策桩和样本的效率较低,若数据集较大,可考虑用向量化操作替代循环,但小数据集下无需调整。

内容的提问来源于stack exchange,提问作者Gio Formichella

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 09:05:57