机器学习中验证集的使用时机及代码集成方法咨询
Hey there! Let's break down how to fix and integrate your validation set into this decision tree regression workflow. You've got the right idea with splitting your data, but there are a few critical adjustments needed to avoid data leakage and use the validation set effectively for hyperparameter tuning.
First: Fix the Scaling Step (Avoid Data Leakage!)
Your current code uses fit_transform() on the test set, which is a mistake. The scaler should only be fit on the training set—using fit() on the test/validation set would leak information about unseen data into your preprocessing, leading to overly optimistic performance estimates. Instead, use transform() on both validation and test sets with the scaler fitted on training data.
Second: Use Validation Set for Hyperparameter Tuning
Your current hyperparameter loop trains models on the entire dataset and evaluates on the test set, which defeats the purpose of the validation set. The validation set exists to test different hyperparameters before touching the test set. Here's the correct flow:
- Train models on the training set with different
max_depthvalues - Evaluate each model on the validation set to measure generalization performance
- Pick the
max_depththat gives the best validation set performance - Retrain a final model on the combined training + validation set (optional but recommended)
- Evaluate this final model on the test set for your final performance metric
Revised Full Code
import numpy as np import matplotlib.pyplot as plt from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler, MinMaxScaler, RobustScaler from sklearn.tree import DecisionTreeRegressor from sklearn.metrics import mean_squared_error # Split the set into train, validation, and test set (70:15:15 for train:valid:test) X_train, X_rem, y_train, y_rem = train_test_split(X, y, train_size=0.7, random_state=42) # Add random_state for reproducibility X_valid, X_test, y_valid, y_test = train_test_split(X_rem, y_rem, test_size=0.5, random_state=42) print("Properties (shapes):\nTraining set: {}\nValidation set: {}\nTest set: {}".format(X_train.shape, X_valid.shape, X_test.shape)) import warnings # supress warnings warnings.filterwarnings('ignore') # SCALING - Correct approach: fit only on training data std = StandardScaler() minmax = MinMaxScaler() rob = RobustScaler() # Transforming the TRAINING set X_train_Standard = std.fit_transform(X_train) X_train_MinMax = minmax.fit_transform(X_train) X_train_Robust = rob.fit_transform(X_train) # Transforming the VALIDATION set (use transform, NOT fit_transform!) X_valid_Standard = std.transform(X_valid) X_valid_MinMax = minmax.transform(X_valid) X_valid_Robust = rob.transform(X_valid) # Transforming the TEST set (same as validation) X_test_Standard = std.transform(X_test) X_test_MinMax = minmax.transform(X_test) X_test_Robust = rob.transform(X_test) # Test scalers for decision tree regressor (now include validation set scores!) treeStd = DecisionTreeRegressor(max_depth=3, random_state=0).fit(X_train_Standard, y_train) treeMinMax = DecisionTreeRegressor(max_depth=3, random_state=0).fit(X_train_MinMax, y_train) treeRobust = DecisionTreeRegressor(max_depth=3, random_state=0).fit(X_train_Robust, y_train) print("Decision tree with standard scaler:\nTraining set score: {:.4f}\nValidation set score: {:.4f}\nTest set score: {:.4f}\n".format( treeStd.score(X_train_Standard, y_train), treeStd.score(X_valid_Standard, y_valid), treeStd.score(X_test_Standard, y_test) )) print("Decision tree with min/max scaler:\nTraining set score: {:.4f}\nValidation set score: {:.4f}\nTest set score: {:.4f}\n".format( treeMinMax.score(X_train_MinMax, y_train), treeMinMax.score(X_valid_MinMax, y_valid), treeMinMax.score(X_test_MinMax, y_test) )) print("Decision tree with robust scaler:\nTraining set score: {:.4f}\nValidation set score: {:.4f}\nTest set score: {:.4f}\n".format( treeRobust.score(X_train_Robust, y_train), treeRobust.score(X_valid_Robust, y_valid), treeRobust.score(X_test_Robust, y_test) )) # Hyperparameter tuning with VALIDATION set (not test set!) max_depths = range(1, 30) training_error = [] validation_error = [] # Replace testing_error with validation_error for max_depth in max_depths: # Train ONLY on training set model = DecisionTreeRegressor(max_depth=max_depth, random_state=42) model.fit(X_train_Standard, y_train) # Using standard scaler as example; repeat for others if needed # Record training error y_pred_train = model.predict(X_train_Standard) training_error.append(mean_squared_error(y_train, y_pred_train)) # Record validation error y_pred_valid = model.predict(X_valid_Standard) validation_error.append(mean_squared_error(y_valid, y_pred_valid)) # Plot training vs validation error plt.figure(figsize=(10,6)) plt.plot(max_depths, training_error, color='blue', label='Training error') plt.plot(max_depths, validation_error, color='green', label='Validation error') plt.xlabel('Tree depth') # Find the optimal max_depth (lowest validation error) optimal_depth = max_depths[np.argmin(validation_error)] plt.axvline(x=optimal_depth, color='orange', linestyle='--') plt.annotate(f'Optimum depth = {optimal_depth}', xy=(optimal_depth-5, min(validation_error)*1.1), color='red') plt.ylabel('Mean squared error') plt.title('Hyperparameter Tuning: Max Depth', pad=20, size=16) plt.legend() plt.show() # Train final model on combined training + validation set (optional but better for data usage) X_train_valid = np.concatenate([X_train_Standard, X_valid_Standard]) y_train_valid = np.concatenate([y_train, y_valid]) final_model = DecisionTreeRegressor(max_depth=optimal_depth, random_state=42) final_model.fit(X_train_valid, y_train_valid) # Final evaluation on TEST set (only do this once!) test_mse = mean_squared_error(y_test, final_model.predict(X_test_Standard)) print(f"Final model test set MSE: {test_mse:.4f}")
Key Explanations
- Scaling Fix: By fitting scalers only on the training set, you ensure your model never sees validation/test data during preprocessing, which keeps your performance estimates honest.
- Validation Set for Tuning: Using the validation set to evaluate hyperparameters means your test set remains completely untouched until the final step—this gives you a true measure of how your model will perform on unseen data.
- Final Model Training: Combining training and validation sets for the final model uses all available data (excluding test) to build the strongest possible model before the final evaluation.
- Reproducibility: Adding
random_statetotrain_test_splitand model initialization ensures your results are consistent every time you run the code.
内容的提问来源于stack exchange,提问作者Mampenda

