不平衡数据集下SVM、随机森林与逻辑回归模型的交叉验证训练及网格搜索参数调优问题咨询
Hey there! Let's break down how to tackle your two core tasks—training models using all cross-validation folds and implementing grid search for hyperparameter tuning—without getting stuck in messy manual loops. Since you're working with an imbalanced dataset, we’ll keep stratification front and center every step of the way.
1. Training & Evaluating Models Using All StratifiedKFold Folds
First, let’s clarify two key goals here: evaluating model performance across all folds to compare candidates, and training a final model on the full training dataset once you’ve picked a winner.
Evaluate Model Performance Across All Folds
You don’t need to write manual for loops for this—scikit-learn’s cross_validate handles the StratifiedKFold iteration automatically, and returns consistent, stratified performance metrics. This is perfect for comparing your three baseline models:
from sklearn.model_selection import cross_validate from sklearn.svm import SVC from sklearn.ensemble import RandomForestClassifier from sklearn.linear_model import LogisticRegression from sklearn.metrics import f1_score, recall_score # Initialize models with class weight adjustment for imbalance models = { "Logistic Regression": LogisticRegression(random_state=42, max_iter=1000, class_weight="balanced"), "Random Forest": RandomForestClassifier(random_state=42, class_weight="balanced"), "SVM": SVC(random_state=42, class_weight="balanced") } # Use StratifiedKFold you defined earlier (skf) for name, model in models.items(): # Track metrics critical for imbalanced data (F1, recall > accuracy) cv_results = cross_validate( model, X_train, y_train, cv=skf, scoring=["accuracy", "f1", "recall"], return_train_score=True ) print(f"--- {name} Cross-Validation Results ---") print(f"Mean Test F1: {cv_results['test_f1'].mean():.4f}") print(f"Mean Test Recall: {cv_results['test_recall'].mean():.4f}\n")
This runs the model through all 10 stratified folds, trains on each fold’s training split, evaluates on the validation split, and averages the results. You’ll get a clear picture of which model performs best on your imbalanced data.
Train a Final Model on the Full Training Set
Once you’ve identified a top-performing model (e.g., Random Forest), train it on the entire training dataset (all folds combined) for your final deployment-ready model:
# Pick your top model from the cross-validation results final_model = RandomForestClassifier(random_state=42, class_weight="balanced") final_model.fit(X_train, y_train) # Evaluate on held-out test set test_preds = final_model.predict(X_test) print(f"Final Model Test F1: {f1_score(y_test, test_preds):.4f}") print(f"Final Model Test Recall: {recall_score(y_test, test_preds):.4f}")
2. Grid Search for Optimal Hyperparameters
Again, no manual double loops needed—GridSearchCV combines grid search and stratified cross-validation into one tool. It will test every parameter combination across all your folds, then return the model with the best average performance.
Step 1: Define Parameter Grids for Each Model
Tailor grids to each model’s key hyperparameters:
param_grids = { "Logistic Regression": { "C": [0.001, 0.01, 0.1, 1, 10, 100], "penalty": ["l1", "l2"], "solver": ["liblinear"] # Required for L1 regularization }, "Random Forest": { "n_estimators": [50, 100, 200], "max_depth": [None, 10, 20, 30], "min_samples_split": [2, 5, 10] }, "SVM": { "C": [0.001, 0.01, 0.1, 1, 10], "kernel": ["linear", "rbf"], "gamma": ["scale", "auto"] } }
Step 2: Run Grid Search for Each Model
Use F1 score as the primary metric (ideal for imbalanced data) to rank parameter combinations:
from sklearn.model_selection import GridSearchCV from sklearn.metrics import make_scorer # Use F1 as the scoring metric for grid search scorer = make_scorer(f1_score) best_trained_models = {} for name, model in models.items(): grid_search = GridSearchCV( estimator=model, param_grid=param_grids[name], cv=skf, scoring=scorer, n_jobs=-1, # Use all CPU cores for speed verbose=1 ) grid_search.fit(X_train, y_train) best_trained_models[name] = grid_search.best_estimator_ print(f"--- Best {name} ---") print(f"Optimal Parameters: {grid_search.best_params_}") print(f"Best Cross-Validation F1: {grid_search.best_score_:.4f}\n")
Step 3: Compare Tuned Models on the Test Set
Finally, evaluate all tuned models on your held-out test set to pick the absolute winner:
for name, model in best_trained_models.items(): test_preds = model.predict(X_test) print(f"--- {name} Tuned Test Performance ---") print(f"Test F1: {f1_score(y_test, test_preds):.4f}") print(f"Test Recall: {recall_score(y_test, test_preds):.4f}\n")
Do You Need Double Loops?
No! Scikit-learn’s built-in tools (cross_validate, GridSearchCV) already handle the nested cross-validation loops internally. Manual loops are error-prone (your original code had a mix-up with X vs X_train indices) and less efficient. Stick to the built-ins for clean, reliable code.
内容的提问来源于stack exchange,提问作者Gvasiles

