如何遍历Sklearn RandomizedSearchCV中的所有随机模型并评估其精度与预测速度?
Great question! When working with RandomizedSearchCV, you can’t directly iterate over the search object to access all trained models—here’s how to properly pull all tested parameter configurations, evaluate their test-set performance, and capture both accuracy and inference speed:
Method 1: Use cv_results_ (Recommended)
The cv_results_ attribute of RandomizedSearchCV stores all metadata from your search, including every parameter combination tested, cross-validation scores, and more. While the search doesn’t save full trained model instances for every configuration, you can reuse the parameters to retrain models on your full training set and evaluate them on the test set:
import pandas as pd import time from sklearn.metrics import mean_squared_error from sklearn.ensemble import RandomForestRegressor # Initialize an empty list to store results results = [] # Iterate through each parameter combo and its CV score from the search for params, cv_mean_score in zip(rf_random.cv_results_['params'], rf_random.cv_results_['mean_test_score']): # Create a model with the current parameter set rf_model = RandomForestRegressor(**params) # Retrain on the full training set (since CV only trains on folds) rf_model.fit(X_train, y_train) # Measure inference speed per sample start_time = time.time() y_pred = rf_model.predict(X_test) infer_time_per_sample = (time.time() - start_time) / X_test.shape[0] # Calculate test-set RMSE rmse = mean_squared_error(y_test, y_pred, squared=False) # Store results (expand with more metrics if needed) results.append({ 'parameters': params, 'test_rmse': rmse, 'inference_time_per_sample': infer_time_per_sample, 'cv_mean_score': cv_mean_score })
Getting Model Parameters
- If you need the full set of model parameters (including defaults not in your grid), use
rf_model.get_params()instead of theparamsvariable fromcv_results_. - The
paramsfromcv_results_only includes the parameters you defined in your search grid, which is often sufficient for tradeoff analysis.
Method 2: Custom Callback (Advanced)
If you want to capture results during the search process without retraining, you can implement a custom callback with sklearn’s BaseSearchCV. However, this requires more boilerplate code, and cv_results_ is almost always the simpler, more reliable choice for most use cases.
Alternative Approaches to Find Accuracy-Speed Balance
Your random search approach is solid, but here are two alternatives that might help you find the optimal balance more efficiently:
Halving Random Search (
HalvingRandomSearchCV)
This sklearn class iteratively eliminates poor-performing parameter combinations, focusing resources on the most promising ones. It’s faster than standard random search, especially for large datasets, as it reduces the number of full model trainings needed.Bayesian Optimization
Libraries likescikit-optimizeofferBayesSearchCV, which uses Bayesian methods to select parameter combinations based on previous results. This is more efficient than random search because it learns from past trials to prioritize configurations that are likely to improve your accuracy-speed tradeoff.
内容的提问来源于stack exchange,提问作者Carlo

