基于scikit-learn的高斯过程回归留一法测试与3D散点绘图
Hey there! Let's build out this Gaussian Process Regression (GPR) workflow with Leave-One-Out (LOO) cross-validation, test multiple kernels, and create those 3D scatter plots for your dataset. I'll break this down into actionable steps with full code you can run right away.
First, let's grab the packages we need for data handling, modeling, cross-validation, and plotting:
import numpy as np from sklearn.gaussian_process import GaussianProcessRegressor from sklearn.gaussian_process.kernels import RBF, Matern, RationalQuadratic, WhiteKernel from sklearn.model_selection import LeaveOneOut import matplotlib.pyplot as plt from mpl_toolkits.mplot3d import Axes3D
Here's your input and target data formatted as numpy arrays (exact as you provided):
# Input data (5 samples, 3 features) inputData = np.array([ [30.1678, -173.569, 725.724], [29.9895, -173.34, 725.76 ], [29.9411, -173.111, 725.768], [29.9306, -173.016, 725.98 ], [29.6754, -172.621, 725.795] ]) # Target data (5 samples, 3 output dimensions) targetData = np.array([ [14.8016, -175.911, 779.752], [14.7319, -175.483, 779.504], [14.5022, -175.087, 779.388], [14.4904, -174.576, 779.416], [14.4881, -174.058, 779.452] ])
We'll test three common kernels (RBF, Matern, RationalQuadratic) plus a WhiteKernel to account for noise. For each LOO fold, we train on 4 samples and predict on the left-out one, storing all results for later analysis.
# Define kernels to test (each paired with noise-handling WhiteKernel) kernels = { "RBF": RBF() + WhiteKernel(), "Matern": Matern(nu=1.5) + WhiteKernel(), "RationalQuadratic": RationalQuadratic() + WhiteKernel() } # Initialize storage for true vs predicted values per kernel loo_results = {kernel: {"true": [], "pred": []} for kernel in kernels.keys()} # Set up Leave-One-Out splitter loo = LeaveOneOut() # Iterate over each LOO fold for train_idx, test_idx in loo.split(inputData): # Split data into train/test for this fold X_train, X_test = inputData[train_idx], inputData[test_idx] y_train, y_test = targetData[train_idx], targetData[test_idx] # Test each kernel for kernel_name, kernel in kernels.items(): # Train GPR model gpr = GaussianProcessRegressor(kernel=kernel, random_state=42) gpr.fit(X_train, y_train) # Predict on the held-out sample y_pred, _ = gpr.predict(X_test, return_std=True) # Store results loo_results[kernel_name]["true"].append(y_test[0]) loo_results[kernel_name]["pred"].append(y_pred[0]) # Convert lists to numpy arrays for easier manipulation for kernel_name in loo_results.keys(): loo_results[kernel_name]["true"] = np.array(loo_results[kernel_name]["true"]) loo_results[kernel_name]["pred"] = np.array(loo_results[kernel_name]["pred"])
We'll make a separate 3D plot for each kernel, comparing true target values (blue) to predicted values (red) across all three output columns.
# Set up figure with 3 subplots (one per kernel) fig = plt.figure(figsize=(18, 5)) # Iterate over kernels and generate plots for idx, (kernel_name, results) in enumerate(loo_results.items(), 1): ax = fig.add_subplot(1, 3, idx, projection='3d') # Extract true and predicted values for each column true_cols = results["true"].T pred_cols = results["pred"].T # Plot true and predicted points ax.scatter(*true_cols, c='blue', label='True Values') ax.scatter(*pred_cols, c='red', marker='^', label='Predicted Values') # Add labels and title ax.set_xlabel('Target Column 1') ax.set_ylabel('Target Column 2') ax.set_zlabel('Target Column 3') ax.set_title(f'GPR with {kernel_name} Kernel (LOO Results)') ax.legend() plt.tight_layout() plt.show()
To quantify performance, let's calculate the Mean Absolute Error (MAE) for each kernel across all output columns:
for kernel_name, results in loo_results.items(): mae_per_col = np.mean(np.abs(results["true"] - results["pred"]), axis=0) avg_mae = np.mean(mae_per_col) print(f"Kernel: {kernel_name}") print(f"MAE per target column: {mae_per_col.round(4)}") print(f"Average MAE across columns: {avg_mae.round(4)}\n")
Quick Notes:
- Kernel Selection: Matern works well for non-smooth data, RBF is ideal for smooth functions, and RationalQuadratic handles data with multiple length scales.
- WhiteKernel: This adds noise modeling, which helps GPR generalize better to small datasets like yours.
- LOO Fit: With only 5 samples, LOO is perfect—it uses nearly all your data for training each iteration, maximizing model quality per fold.
内容的提问来源于stack exchange,提问作者nasil122002

