You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AdaBoostClassifier无estimator_属性报错:单轮树构建计时遇问题

AdaBoost每轮决策树构建计时问题解决

问题背景

需要统计AdaBoost算法每一轮新增决策树的构建时长,使用scikit-learn 1.4.0版本,现有代码运行时抛出AttributeError: 'AdaBoostClassifier' object has no attribute 'estimator_'错误。

原代码:

Y, z = parse.getHARData() #returns my features Y and labels z

Z_train, Z_test, j_train, j_test = train_test_split(Y, z, test_size=0.30, shuffle=True)

b_estimator = DecisionTreeClassifier(max_depth=DEPTH)

ada = AdaBoostClassifier(estimator=b_estimator, n_estimators=NUMTREES)

elapsed_times = []

for stage in range(NUMTREES): 
    start_time = time.time()

    # Access and fit the current base estimator
    base_estimator = ada._make_estimator(append=True, random_state=42)
    base_estimator.fit(Z_train, j_train)
    
    elapsed_time = time.time() - start_time
    elapsed_times.append(elapsed_time)

错误栈:

AttributeError                            Traceback (most recent call last)
Cell In[9], line 7
      4 start_time = time.time()
      6 # Access and fit the current base estimator
----> 7 base_estimatr = ada._make_estimator(append=True, random_state=42)
      8 base_estimatr.fit(Z_train, j_train)
     10 elapsed_time = time.time() - start_time

File ~/anaconda3/envs/ADA/lib/python3.11/site-packages/sklearn/ensemble/_base.py:141, in BaseEnsemble._make_estimator(self, append, random_state)
    135 def _make_estimator(self, append=True, random_state=None):
    136     """Make and configure a copy of the `estimator_` attribute.
    137 
    138     Warning: This method should be used to properly instantiate new
    139     sub-estimators.
    140     """
--> 141     estimator = clone(self.estimator_)
    142     estimator.set_params(**{p: getattr(self, p) for p in self.estimator_params})
    144     if random_state is not None:

AttributeError: 'AdaBoostClassifier' object has no attribute 'estimator_'

错误原因

  1. _make_estimator是scikit-learn内部私有方法,依赖estimator_属性,但该属性仅在模型完成首次拟合后才会被初始化,直接调用会因属性未创建报错。
  2. 原代码手动调用base_estimator.fit()没有利用AdaBoost的核心逻辑(样本权重更新、基模型权重计算),相当于训练了NUMTREES个独立的决策树,完全脱离了AdaBoost集成机制。

修正方案

手动实现AdaBoost的逐轮训练流程,同时记录每一轮基模型的构建时长:

import time
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.base import clone
import numpy as np

# 加载数据
Y, z = parse.getHARData()
Z_train, Z_test, j_train, j_test = train_test_split(Y, z, test_size=0.30, shuffle=True)

# 初始化参数
DEPTH = 1  # 根据你的需求调整
NUMTREES = 10  # 根据你的需求调整
base_estimator = DecisionTreeClassifier(max_depth=DEPTH)

# 初始化AdaBoost相关变量
n_samples = Z_train.shape[0]
sample_weights = np.ones(n_samples) / n_samples
estimators = []
estimator_weights = []
elapsed_times = []

for _ in range(NUMTREES):
    start_time = time.time()
    
    # 克隆基估计器并拟合加权样本
    estimator = clone(base_estimator)
    estimator.fit(Z_train, j_train, sample_weight=sample_weights)
    
    # 计算错误率和基模型权重
    y_pred = estimator.predict(Z_train)
    incorrect = y_pred != j_train
    error_rate = np.mean(np.average(incorrect, weights=sample_weights, axis=0))
    
    # 避免除零,设置最小误差阈值
    error_rate = max(error_rate, 1e-10)
    alpha = np.log((1 - error_rate) / error_rate)
    
    # 更新样本权重
    sample_weights *= np.exp(alpha * incorrect)
    sample_weights /= sample_weights.sum()
    
    # 计时结束并记录
    elapsed_time = time.time() - start_time
    elapsed_times.append(elapsed_time)
    
    # 保存基模型和权重
    estimators.append(estimator)
    estimator_weights.append(alpha)

# 可选:将训练好的模型组装成AdaBoostClassifier(如果需要后续使用)
ada = AdaBoostClassifier(estimator=base_estimator, n_estimators=NUMTREES)
ada.estimators_ = estimators
ada.estimator_weights_ = estimator_weights
ada.classes_ = np.unique(j_train)

说明

  • 该方案严格遵循AdaBoost的训练逻辑:每一轮根据样本权重拟合基模型,计算模型权重并更新样本权重,同时准确记录单轮模型构建的耗时。
  • 最后可选将训练好的基模型和权重赋值给AdaBoostClassifier实例,方便后续进行预测等操作。

内容的提问来源于stack exchange,提问作者Joed Ngangmeni

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 09:44:56