You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

不平衡数据集下SVM、随机森林与逻辑回归模型的交叉验证训练及网格搜索参数调优问题咨询

Hey there! Let's break down how to tackle your two core tasks—training models using all cross-validation folds and implementing grid search for hyperparameter tuning—without getting stuck in messy manual loops. Since you're working with an imbalanced dataset, we’ll keep stratification front and center every step of the way.


1. Training & Evaluating Models Using All StratifiedKFold Folds

First, let’s clarify two key goals here: evaluating model performance across all folds to compare candidates, and training a final model on the full training dataset once you’ve picked a winner.

Evaluate Model Performance Across All Folds

You don’t need to write manual for loops for this—scikit-learn’s cross_validate handles the StratifiedKFold iteration automatically, and returns consistent, stratified performance metrics. This is perfect for comparing your three baseline models:

from sklearn.model_selection import cross_validate
from sklearn.svm import SVC
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import f1_score, recall_score

# Initialize models with class weight adjustment for imbalance
models = {
    "Logistic Regression": LogisticRegression(random_state=42, max_iter=1000, class_weight="balanced"),
    "Random Forest": RandomForestClassifier(random_state=42, class_weight="balanced"),
    "SVM": SVC(random_state=42, class_weight="balanced")
}

# Use StratifiedKFold you defined earlier (skf)
for name, model in models.items():
    # Track metrics critical for imbalanced data (F1, recall > accuracy)
    cv_results = cross_validate(
        model, X_train, y_train, cv=skf,
        scoring=["accuracy", "f1", "recall"],
        return_train_score=True
    )
    print(f"--- {name} Cross-Validation Results ---")
    print(f"Mean Test F1: {cv_results['test_f1'].mean():.4f}")
    print(f"Mean Test Recall: {cv_results['test_recall'].mean():.4f}\n")

This runs the model through all 10 stratified folds, trains on each fold’s training split, evaluates on the validation split, and averages the results. You’ll get a clear picture of which model performs best on your imbalanced data.

Train a Final Model on the Full Training Set

Once you’ve identified a top-performing model (e.g., Random Forest), train it on the entire training dataset (all folds combined) for your final deployment-ready model:

# Pick your top model from the cross-validation results
final_model = RandomForestClassifier(random_state=42, class_weight="balanced")
final_model.fit(X_train, y_train)

# Evaluate on held-out test set
test_preds = final_model.predict(X_test)
print(f"Final Model Test F1: {f1_score(y_test, test_preds):.4f}")
print(f"Final Model Test Recall: {recall_score(y_test, test_preds):.4f}")

2. Grid Search for Optimal Hyperparameters

Again, no manual double loops needed—GridSearchCV combines grid search and stratified cross-validation into one tool. It will test every parameter combination across all your folds, then return the model with the best average performance.

Step 1: Define Parameter Grids for Each Model

Tailor grids to each model’s key hyperparameters:

param_grids = {
    "Logistic Regression": {
        "C": [0.001, 0.01, 0.1, 1, 10, 100],
        "penalty": ["l1", "l2"],
        "solver": ["liblinear"]  # Required for L1 regularization
    },
    "Random Forest": {
        "n_estimators": [50, 100, 200],
        "max_depth": [None, 10, 20, 30],
        "min_samples_split": [2, 5, 10]
    },
    "SVM": {
        "C": [0.001, 0.01, 0.1, 1, 10],
        "kernel": ["linear", "rbf"],
        "gamma": ["scale", "auto"]
    }
}

Step 2: Run Grid Search for Each Model

Use F1 score as the primary metric (ideal for imbalanced data) to rank parameter combinations:

from sklearn.model_selection import GridSearchCV
from sklearn.metrics import make_scorer

# Use F1 as the scoring metric for grid search
scorer = make_scorer(f1_score)

best_trained_models = {}

for name, model in models.items():
    grid_search = GridSearchCV(
        estimator=model,
        param_grid=param_grids[name],
        cv=skf,
        scoring=scorer,
        n_jobs=-1,  # Use all CPU cores for speed
        verbose=1
    )
    grid_search.fit(X_train, y_train)
    
    best_trained_models[name] = grid_search.best_estimator_
    print(f"--- Best {name} ---")
    print(f"Optimal Parameters: {grid_search.best_params_}")
    print(f"Best Cross-Validation F1: {grid_search.best_score_:.4f}\n")

Step 3: Compare Tuned Models on the Test Set

Finally, evaluate all tuned models on your held-out test set to pick the absolute winner:

for name, model in best_trained_models.items():
    test_preds = model.predict(X_test)
    print(f"--- {name} Tuned Test Performance ---")
    print(f"Test F1: {f1_score(y_test, test_preds):.4f}")
    print(f"Test Recall: {recall_score(y_test, test_preds):.4f}\n")

Do You Need Double Loops?

No! Scikit-learn’s built-in tools (cross_validate, GridSearchCV) already handle the nested cross-validation loops internally. Manual loops are error-prone (your original code had a mix-up with X vs X_train indices) and less efficient. Stick to the built-ins for clean, reliable code.

内容的提问来源于stack exchange,提问作者Gvasiles

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 06:07:34