自定义Sklearn AdaBoostClassifier报错:fit未返回self触发AssertionError
嘿,这个问题的根源其实特别明显——你看报错里的断言assert estimator.fit(X, y) is estimator,scikit-learn要求所有估算器的fit方法必须返回自身实例(也就是self),但你的AdaBoostClassifier的fit方法最后返回的是self.models列表,这就直接触发了这个检查失败的断言错误。
scikit-learn的这个规范是为了保证API的一致性,比如支持clf.fit(X,y).predict(X)这种链式调用,同时也是check_estimator验证的强制要求之一。
修正后的核心代码
只需要修改fit方法的最后一行,把返回self.models改成返回self就行。另外我还顺便优化了一些细节,让你的实现更符合AdaBoost的设计和sklearn的规范:
from sklearn.base import ClassifierMixin, BaseEstimator, clone from sklearn.utils.validation import check_X_y, check_array, check_is_fitted from sklearn.utils.multiclass import unique_labels from sklearn.tree import DecisionTreeClassifier from sklearn.metrics import accuracy_score import numpy as np class AdaBoostClassifier(ClassifierMixin, BaseEstimator): def __init__(self, base_estimator=None, n_estimators=50, random_state=None): self.base_estimator = base_estimator self.n_estimators = n_estimators self.random_state = random_state def fit(self, X, y): """ ---------- X : array-like, shape (n_samples, n_features) The training input samples. y : array-like, shape (n_samples,) The target values. An array of int. Returns ------- self : object Returns self. """ # Check that X and y have correct shape X, y = check_X_y(X, y) # Store the classes seen during fit self.classes_ = unique_labels(y) self.X_ = X self.y_ = y self.models = [] self.alphas = [] n_samples, _ = X.shape w = np.ones(n_samples) / n_samples # 使用初始化时传入的基础估算器,默认用深度为1的决策树 base_clf = self.base_estimator or DecisionTreeClassifier(max_depth=1, random_state=self.random_state) for m in range(self.n_estimators): # 克隆基础估算器,避免多个模型共享同一实例状态 clf = clone(base_clf) clf.fit(X, y, sample_weight=w) pred = clf.predict(X) error = w.dot(pred != y) # 防止error为0导致log报错 error = max(error, 1e-10) alpha = 0.5 * (np.log(1 - error) - np.log(error)) w = w * np.exp(-alpha * y * pred) w = w / w.sum() # 归一化权重 self.models.append(clf) self.alphas.append(alpha) # 关键修正:返回self而非models列表 return self def predict(self, X): """ A reference implementation of a prediction for a classifier. Parameters ---------- X : array-like, shape (n_samples, n_features) The input samples. Returns ------- y : ndarray, shape (n_samples,) The predicted labels for each sample. """ # 确保fit方法已被调用,检查必要的属性 check_is_fitted(self, ['X_', 'y_', 'models', 'alphas']) # 输入验证 X = check_array(X) n_samples, _ = X.shape ada_scores = np.zeros(n_samples) for alpha, clf in zip(self.alphas, self.models): ada_scores += alpha * clf.predict(X) # 取符号得到最终预测结果 return np.sign(ada_scores) def score(self, X, y): pred = self.predict(X) # 使用sklearn内置的accuracy_score,更符合规范 return accuracy_score(y, pred)
额外的优化点说明
- 启用base_estimator参数:原来的代码没有用到初始化时传入的
base_estimator,现在改成如果用户传入了基础模型就用它,否则默认用深度为1的决策树,更符合AdaBoost的设计。 - 克隆基础模型:每次迭代都用
clone(base_clf)创建新的模型实例,避免多个模型共享同一个实例的状态,防止意外的副作用。 - 处理极端误差值:加入
error = max(error, 1e-10),防止模型完全拟合训练数据时error为0,导致np.log(error)抛出除以0的错误。 - 完善fit检查:在
predict方法里的check_is_fitted加入了models和alphas,确保这些必要属性已经被初始化。 - 规范score方法:改用sklearn内置的
accuracy_score计算准确率,符合sklearn的API标准。
内容的提问来源于stack exchange,提问作者Ryantstrong
相关产品推荐
相关产品推荐

