为何GridSearchCV在同一数据集上未返回最优得分?
Let's break down what's happening here—you've run into a common pitfall with combining GridSearchCV and Pipeline, plus some confusion about how your models are actually being trained.
The Core Issue: Duplicate Fitting & Data Scaling Mismatch
Looking at both code blocks, you're making the same critical mistake:
- First, you fit the
GridSearchCVobject directly on your raw data (X). - Then, you wrap that already-fit
GridSearchCVinto aPipelinewithRobustScalerand fit the pipeline again.
This does two things that mess up your results:
- The second fit call on the pipeline overwrites all the results (
best_score_,best_params_) from the first fit. - The first fit uses raw, unscaled data, while the second fit uses data transformed by
RobustScaler. The scores you're comparing are from two completely different datasets!
Let's Walk Through Your Code
First Code Block
- When you run
grid.fit(X,train_labels),GridSearchCVruns cross-validation on your rawXand stores results. - Then
gridCV.fit(X,train_labels)scalesXfirst, then runsGridSearchCVagain on the scaled data. This overwrites the originalgridobject's results. - The
best_score_andbest_params_you see are from the scaled data, where the combination{'C': 25, 'epsilon': 0.005, 'gamma': 0.0003}truly has the highest score (0.90146).
Second Code Block
- Same pattern: you first fit
GridSearchCVon raw data, then overwrite it by fitting the pipeline on scaled data. - The score 0.90103 is for the single parameter set on the scaled data, which is lower than the best score from the full grid search. That's why
GridSearchCVdidn't pick it—it wasn't the best option on the data it was actually evaluated on.
The Fix: Proper Pipeline Usage
To avoid data leakage and get accurate results, you should never fit GridSearchCV separately before putting it in a pipeline. The pipeline should handle scaling within each cross-validation fold (so validation data never sees the scaling parameters from the training data). Here's the correct approach:
from sklearn.pipeline import make_pipeline from sklearn.preprocessing import RobustScaler from sklearn.svm import SVR from sklearn.model_selection import GridSearchCV # Define your parameter grid params = {'C':[15,20,25],'epsilon':[.003,.005,.008],'gamma':[.0001,.0003,.0006]} # Create pipeline with scaler and GridSearchCV (no pre-fitting!) gridCV = make_pipeline(RobustScaler(), GridSearchCV(SVR(), param_grid=params)) # Fit once—this handles scaling *per fold* during cross-validation gridCV.fit(X, train_labels) # Extract the GridSearchCV results from the pipeline grid_search_results = gridCV.named_steps['gridsearchcv'] print("Best score:", grid_search_results.best_score_) print("Best parameters:", grid_search_results.best_params_)
Key Takeaway
Your original results were skewed because you overwrote the GridSearchCV object with a second fit on scaled data. The full grid search did find the true best parameter combination for the scaled dataset, which outperformed the single parameter set you tested later.
内容的提问来源于stack exchange,提问作者ankur.salunke

