能否将OneVsRestClassifier设为AdaBoost基估计器及网格搜索参数配置咨询
Let's break down your issues step by step, fix the code, and add optimizations tailored to your binary classification use case:
First: Fix Immediate Syntax Errors
Your code has two obvious syntax bugs that are causing failures:
- The
AdaBoostClassifierinitialization line is missing a closing parenthesis at the end. - Your
printstatements use incorrect formatting syntax—format()is a method of the string object, so it needs to be called directly on your output string.
Second: Optimize Your Approach (Binary Targets Don't Need OneVsRest)
Since you've already binarized your target variables (yb_train2 and yb_train3 are binary labels), OneVsRestClassifier is unnecessary—it's designed to adapt binary classifiers for multi-class tasks. Using it here adds unnecessary complexity to your parameter grid. Instead, you can directly use DecisionTreeClassifier as the base estimator for AdaBoost.
Corrected Code (Recommended for Binary Tasks)
from sklearn.ensemble import AdaBoostClassifier from sklearn.tree import DecisionTreeClassifier from sklearn.model_selection import GridSearchCV # Initialize AdaBoost with a decision tree base estimator abc = AdaBoostClassifier(base_estimator=DecisionTreeClassifier()) # Define parameter grid: combine AdaBoost params and decision tree params param_grid = { # Decision tree parameters (prefixed with base_estimator__) "base_estimator__criterion": ["gini", "entropy"], "base_estimator__splitter": ["best", "random"], # AdaBoost's own parameters "n_estimators": [10, 50, 100], # 1/2 is too small—AdaBoost needs enough weak learners "learning_rate": [0.0001, 0.001, 0.01, 0.1, 1] } # Grid search with ROC AUC as the scoring metric (matches your goal) grid = GridSearchCV(abc, param_grid, scoring="roc_auc", cv=5) # Train and evaluate for age group grid.fit(X_train, yb_train2) print(f'Best score for age group: {grid.best_score_}') print(f'Best parameters: {grid.best_params_}') # Train and evaluate for race group grid.fit(X_train, yb_train3) print(f'Best score for race group: {grid.best_score_}') print(f'Best parameters: {grid.best_params_}')
If You Must Keep OneVsRest (For Future Multi-Class Compatibility)
If you want to retain OneVsRestClassifier (e.g., to easily switch back to multi-class labels later), here's the corrected version with proper nested parameter paths:
from sklearn.ensemble import AdaBoostClassifier from sklearn.tree import DecisionTreeClassifier from sklearn.multiclass import OneVsRestClassifier from sklearn.model_selection import GridSearchCV # Fix missing closing parenthesis abc = AdaBoostClassifier(base_estimator=OneVsRestClassifier(DecisionTreeClassifier())) # Parameter grid with nested path: base_estimator (OneVsRest) -> estimator (DecisionTree) param_grid = { "base_estimator__estimator__criterion": ["gini", "entropy"], "base_estimator__estimator__splitter": ["best", "random"], "n_estimators": [10, 50, 100], "learning_rate": [0.0001, 0.001, 0.01, 0.1, 1] } grid = GridSearchCV(abc, param_grid, scoring="roc_auc", cv=5) # Age group training grid.fit(X_train, yb_train2) print(f'Best score for age group: {grid.best_score_}') print(f'Best parameters: {grid.best_params_}') # Race group training grid.fit(X_train, yb_train3) print(f'Best score for race group: {grid.best_score_}') print(f'Best parameters: {grid.best_params_}')
Additional Tips
- Explicit Scoring: Using
scoring="roc_auc"ensures GridSearch uses the exact metric you care about (instead of default accuracy). - n_estimators Tuning: Values like 10, 50, or 100 are more practical than 1/2—AdaBoost relies on combining multiple weak learners to build a strong model.
- Cross-Validation: The
cv=5parameter sets 5-fold cross-validation, making your results more reliable.
内容的提问来源于stack exchange,提问作者Darsolation

