自定义AdaBoost分类器使用LearningCurveDisplay绘制学习曲线报错求助
自定义AdaBoost分类器绘制学习曲线的解决方案
问题背景
为完成学校项目,我开发了一个以决策树桩(DecisionStump)作为弱学习器的AdaBoost分类器,具体实现代码如下:
class AdaBoost: def __init__(self, boosting_rounds): self.boosting_rounds = boosting_rounds # number of weak learners used self.weak_learners = [] # [ (learner, amount of say), ... ] self.decision_stumps = [] def fit(self, x, y): n_rows = x.shape[0] # number of examples in training set attributes = x.columns # features in training set # dataset example weights d = np.full(n_rows, 1 / n_rows) # preparing all decision stumps for a in attributes: # finding possible thresholds for decision stump values = x[a].unique() values.sort() thresholds = [] for i in range(1, len(values)): thresholds.append((values[i] + values[i - 1]) / 2) for threshold in thresholds: self.decision_stumps.append(DecisionStump(a, threshold)) for t in range(0, self.boosting_rounds): # choosing decision stump that minimizes weights dependent error min_error = float("inf") h = None for ds in self.decision_stumps: ds_error = 0 for example in range(0, n_rows): example_row = x.iloc[example] if ds.predict(example_row) != y.iloc[example][0]: # misclassification ds_error += d[example] if ds_error < min_error: min_error = ds_error h = ds # amount of say small = 1e-10 alpha = 0.5 * np.log((1 - min_error) / (min_error + small)) # small avoids division by 0 # storing learner with corresponding amount of say self.weak_learners.append((h, alpha)) # updating weights for i in range(0, len(d)): d[i] = d[i] * np.exp(-alpha * y.iloc[i] * h.predict(x.iloc[i])) z = np.sum(d) # normalisation factor for i in range(0, len(d)): d[i] /= z def predict(self, x): s = 0 # sum of predictions for i in range(0, len(self.weak_learners)): prediction = self.weak_learners[i][0].predict(x) s += prediction * self.weak_learners[i][1] # prediction * amount of say return int(np.sign(s))
尝试使用sklearn.model_selection.LearningCurveDisplay绘制学习曲线时,先是报错提示估算器没有评分函数,添加评分函数后又出现其他问题,需要找到正确的绘制方法。
解决步骤
1. 让自定义AdaBoost符合sklearn估算器规范
sklearn工具类要求估算器要么实现score方法,要么在调用时指定scoring参数。给AdaBoost类添加score方法,计算分类准确率:
from sklearn.metrics import accuracy_score import numpy as np import pandas as pd class AdaBoost: # 保留原有__init__、fit、predict方法... def score(self, X, y): # 处理输入:如果y是DataFrame,转成一维数组 if isinstance(y, pd.DataFrame): y = y.iloc[:, 0].values # 批量预测 predictions = [self.predict(X.iloc[i]) for i in range(X.shape[0])] return accuracy_score(y, predictions)
2. 调整训练数据的y格式
你的fit方法中使用y.iloc[i][0],说明y是DataFrame格式,训练前将y转换为一维数组或Series,避免后续出错:
# 假设原始数据是X_train(DataFrame)、y_train(DataFrame) y_train = y_train.iloc[:, 0].values # 转成一维numpy数组
3. 正确绘制学习曲线
使用LearningCurveDisplay.from_estimator时,确保参数符合要求,示例代码如下:
from sklearn.model_selection import LearningCurveDisplay import matplotlib.pyplot as plt # 初始化自定义AdaBoost分类器 adaboost_clf = AdaBoost(boosting_rounds=10) # 绘制学习曲线 display = LearningCurveDisplay.from_estimator( estimator=adaboost_clf, X=X_train, y=y_train, train_sizes=np.linspace(0.1, 1.0, 5), # 训练集比例 cv=5, # 5折交叉验证 scoring="accuracy", # 显式指定评分指标(可选,已实现score方法时可省略) n_jobs=-1, # 使用所有CPU核心加速 line_kw={"marker": "o"}, random_state=42 ) plt.title("Learning Curve for Custom AdaBoost Classifier") plt.show()
额外问题排查
- 决策树桩兼容性:确保
DecisionStump类的predict方法能正确处理单条样本(如DataFrame的行),返回值为1或-1(与AdaBoost标签格式匹配)。 - 标签格式校验:训练数据的y标签必须是
1和-1,而非0和1,否则会影响权重更新逻辑。 - 性能优化:当前
fit方法遍历所有决策桩和样本的效率较低,若数据集较大,可考虑用向量化操作替代循环,但小数据集下无需调整。
内容的提问来源于stack exchange,提问作者Gio Formichella
相关产品推荐
相关产品推荐

