如何调整GridSearchCV脚本生成各算法最优分类器排序表格
First: What's the Point of That Initial LogisticRegression in Pipeline?
Great question! That initial LogisticRegression() in the Pipeline is just a placeholder estimator. The Pipeline requires you to define the structure of your workflow upfront—so you're telling it "we have a step called classifier that will be some kind of classification model".
When you run GridSearchCV with your search_space, every entry in the search space that uses the classifier key will replace that initial placeholder. You could even swap it out for a DummyClassifier() and the code would work exactly the same, as long as the step name matches what's in your search space. It's just there to satisfy the Pipeline's requirement of having an initial, valid estimator for each step.
Modified Script: Per-Algorithm Best Models + Sorted Metrics Table
Here's a revised script that does exactly what you want: it finds the best model for each algorithm, calculates precision/recall/F-score metrics, and displays them in a clean table sorted by weighted F-score. We'll use pandas to make the table easy to read and sort:
import numpy as np import pandas as pd from sklearn import datasets from sklearn.linear_model import LogisticRegression from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import GridSearchCV from sklearn.metrics import ( accuracy_score, classification_report, precision_recall_fscore_support ) # Load and split data (keeping your original test split for consistency) iris = datasets.load_iris() X_train = iris.data y_train = iris.target X_test = iris.data[:10] y_test = iris.target[:10] seed = 1 # Define models, their base estimators, and parameter grids model_configs = [ { "name": "Logistic Regression", "estimator": LogisticRegression(random_state=seed), "params": { "penalty": ["l1", "l2"], "C": np.logspace(0, 4, 10) } }, { "name": "Random Forest", "estimator": RandomForestClassifier(random_state=seed, n_jobs=-1), "params": { "n_estimators": [10, 100, 1000], "max_features": [1, 2, 3] } } ] # Store results for each model results = [] # Iterate through each model to find its optimal version for config in model_configs: print(f"\n=== Evaluating {config['name']} ===") # Run GridSearchCV for the current model clf = GridSearchCV( estimator=config["estimator"], param_grid=config["params"], scoring="accuracy", refit=True, n_jobs=-1, cv=5, verbose=0 ) clf.fit(X_train, y_train) # Print best training results print(f"Best training accuracy: {clf.best_score_:.3f}") print(f"Best parameters: {clf.best_params_}") # Evaluate on test data y_pred = clf.predict(X_test) test_acc = accuracy_score(y_test, y_pred) print(f"Test set accuracy: {test_acc:.3f}") print(classification_report(y_test, y_pred, digits=4)) # Get weighted precision, recall, f-score, and total support precision, recall, fscore, support = precision_recall_fscore_support( y_test, y_pred, average="weighted" ) # Store results in our list results.append({ "Model Name": config["name"], "Weighted Precision": round(precision, 4), "Weighted Recall": round(recall, 4), "Weighted F-Score": round(fscore, 4), "Total Support": int(support), "Test Accuracy": round(test_acc, 3), "Best Params": clf.best_params_ }) # Convert results to a DataFrame and sort by F-Score (descending) results_df = pd.DataFrame(results).sort_values(by="Weighted F-Score", ascending=False) # Display the sorted metrics table print("\n=== Sorted Model Metrics (by Weighted F-Score) ===") print(results_df.to_string(index=False)) # Highlight the global best model global_best = results_df.iloc[0] print(f"\nGlobal Best Model: {global_best['Model Name']}") print(f"Best F-Score: {global_best['Weighted F-Score']} | Test Accuracy: {global_best['Test Accuracy']}")
Key Features of This Script:
- Per-Algorithm Optima: We loop through each model type separately, so we get the best hyperparameters for Logistic Regression and Random Forest individually.
- Structured Results: All metrics are stored in a list of dictionaries, then converted to a pandas DataFrame for easy sorting and formatting.
- Sorted Table: The table is ordered by weighted F-score (highest first) so you can quickly spot top performers.
- Clear Output: We print detailed results for each model before showing the aggregated table, so you can dig into individual performance if needed.
Why This Approach Works Better Than the Original Pipeline + GridSearch:
The original Pipeline approach is great for finding a single global best model, but if you want to compare the best version of each algorithm side-by-side, looping through each model separately makes it much easier to extract and store results per algorithm. You could adapt the Pipeline method for this, but it would require extra work to split out results for each classifier type from the GridSearchCV results object.
内容的提问来源于stack exchange,提问作者Christopher

