You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义Sklearn AdaBoostClassifier报错:fit未返回self触发AssertionError

嘿,这个问题的根源其实特别明显——你看报错里的断言assert estimator.fit(X, y) is estimator,scikit-learn要求所有估算器的fit方法必须返回自身实例(也就是self),但你的AdaBoostClassifier的fit方法最后返回的是self.models列表,这就直接触发了这个检查失败的断言错误。

scikit-learn的这个规范是为了保证API的一致性,比如支持clf.fit(X,y).predict(X)这种链式调用,同时也是check_estimator验证的强制要求之一。

修正后的核心代码

只需要修改fit方法的最后一行,把返回self.models改成返回self就行。另外我还顺便优化了一些细节,让你的实现更符合AdaBoost的设计和sklearn的规范:

from sklearn.base import ClassifierMixin, BaseEstimator, clone
from sklearn.utils.validation import check_X_y, check_array, check_is_fitted
from sklearn.utils.multiclass import unique_labels
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score
import numpy as np

class AdaBoostClassifier(ClassifierMixin, BaseEstimator):
    def __init__(self, base_estimator=None, n_estimators=50, random_state=None):
        self.base_estimator = base_estimator
        self.n_estimators = n_estimators
        self.random_state = random_state

    def fit(self, X, y):
        """
        ----------
        X : array-like, shape (n_samples, n_features)
        The training input samples.
        y : array-like, shape (n_samples,)
        The target values. An array of int.
        Returns
        -------
        self : object
        Returns self.
        """
        # Check that X and y have correct shape
        X, y = check_X_y(X, y)
        # Store the classes seen during fit
        self.classes_ = unique_labels(y)
        self.X_ = X
        self.y_ = y
        self.models = []
        self.alphas = []
        n_samples, _ = X.shape
        w = np.ones(n_samples) / n_samples
        
        # 使用初始化时传入的基础估算器,默认用深度为1的决策树
        base_clf = self.base_estimator or DecisionTreeClassifier(max_depth=1, random_state=self.random_state)
        
        for m in range(self.n_estimators):
            # 克隆基础估算器,避免多个模型共享同一实例状态
            clf = clone(base_clf)
            clf.fit(X, y, sample_weight=w)
            pred = clf.predict(X)
            error = w.dot(pred != y)
            # 防止error为0导致log报错
            error = max(error, 1e-10)
            alpha = 0.5 * (np.log(1 - error) - np.log(error))
            w = w * np.exp(-alpha * y * pred)
            w = w / w.sum()  # 归一化权重
            self.models.append(clf)
            self.alphas.append(alpha)
        
        # 关键修正:返回self而非models列表
        return self

    def predict(self, X):
        """
        A reference implementation of a prediction for a classifier.
        Parameters
        ----------
        X : array-like, shape (n_samples, n_features)
        The input samples.
        Returns
        -------
        y : ndarray, shape (n_samples,)
        The predicted labels for each sample.
        """
        # 确保fit方法已被调用,检查必要的属性
        check_is_fitted(self, ['X_', 'y_', 'models', 'alphas'])
        # 输入验证
        X = check_array(X)
        n_samples, _ = X.shape
        ada_scores = np.zeros(n_samples)
        for alpha, clf in zip(self.alphas, self.models):
            ada_scores += alpha * clf.predict(X)
        # 取符号得到最终预测结果
        return np.sign(ada_scores)

    def score(self, X, y):
        pred = self.predict(X)
        # 使用sklearn内置的accuracy_score,更符合规范
        return accuracy_score(y, pred)

额外的优化点说明

  1. 启用base_estimator参数:原来的代码没有用到初始化时传入的base_estimator,现在改成如果用户传入了基础模型就用它,否则默认用深度为1的决策树,更符合AdaBoost的设计。
  2. 克隆基础模型:每次迭代都用clone(base_clf)创建新的模型实例,避免多个模型共享同一个实例的状态,防止意外的副作用。
  3. 处理极端误差值:加入error = max(error, 1e-10),防止模型完全拟合训练数据时error为0,导致np.log(error)抛出除以0的错误。
  4. 完善fit检查:在predict方法里的check_is_fitted加入了models和alphas,确保这些必要属性已经被初始化。
  5. 规范score方法:改用sklearn内置的accuracy_score计算准确率,符合sklearn的API标准。

内容的提问来源于stack exchange,提问作者Ryantstrong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 18:47:55