如何联合优化两个scikit-learn模型的特征拆分加性预测拟合?
Great question! The key here is that you don’t want to train the two models separately—because that would optimize each model to predict y on its own, not their sum. Instead, you need to jointly optimize both models by tying their updates to a single loss function that measures how well their combined predictions match y.
Let’s break down how to do this with practical implementations:
Core Idea
Instead of minimizing loss for each model individually (e.g., MSE(f1(x1), y) and MSE(f2(x2), y)), we minimize a single loss that evaluates the sum of their predictions against y:
Loss = MSE(f1(x1) + f2(x2), y)
Both models’ parameters are updated in lockstep based on this shared loss, ensuring they collaborate to fit the additive relationship.
Method 1: Custom Scikit-Learn Estimator (for traditional ML models)
Scikit-learn doesn’t have built-in support for this kind of joint training, but you can create a custom estimator to wrap two models and train them together. Here’s an example using SGDRegressor (works with any model that exposes its parameters):
from sklearn.base import BaseEstimator, RegressorMixin from sklearn.linear_model import SGDRegressor import numpy as np class AdditiveJointModel(BaseEstimator, RegressorMixin): def __init__(self, model1=None, model2=None): # Use SGDRegressor by default, or let users pass custom models self.model1 = model1 if model1 else SGDRegressor(eta0=0.01) self.model2 = model2 if model2 else SGDRegressor(eta0=0.01) def fit(self, X, y, epochs=100, batch_size=32): # Split features into x1 and x2 X1 = X[:, 0].reshape(-1, 1) X2 = X[:, 1].reshape(-1, 1) n_samples = X.shape[0] for _ in range(epochs): # Shuffle data for stochastic training indices = np.random.permutation(n_samples) X1_shuffled, X2_shuffled, y_shuffled = X1[indices], X2[indices], y[indices] # Train in mini-batches for i in range(0, n_samples, batch_size): end = min(i + batch_size, n_samples) x1_batch, x2_batch, y_batch = X1_shuffled[i:end], X2_shuffled[i:end], y_shuffled[i:end] # Compute combined predictions pred1 = self.model1.predict(x1_batch) pred2 = self.model2.predict(x2_batch) total_pred = pred1 + pred2 # Calculate gradient for MSE loss (dLoss/dPred = 2*(total_pred - y_batch)) error = total_pred - y_batch grad1 = 2 * error.reshape(-1, 1) * x1_batch grad2 = 2 * error.reshape(-1, 1) * x2_batch # Update model parameters manually (matches SGD logic) self.model1.coef_ -= self.model1.eta0 * np.mean(grad1, axis=0) self.model1.intercept_ -= self.model1.eta0 * np.mean(2 * error) self.model2.coef_ -= self.model2.eta0 * np.mean(grad2, axis=0) self.model2.intercept_ -= self.model2.eta0 * np.mean(2 * error) return self def predict(self, X): X1 = X[:, 0].reshape(-1, 1) X2 = X[:, 1].reshape(-1, 1) return self.model1.predict(X1) + self.model2.predict(X2)
Method 2: Deep Learning Framework (PyTorch/TensorFlow, more flexible)
For neural networks or complex models, using a deep learning framework simplifies joint optimization—since it handles gradient computation and parameter updates automatically. Here’s a PyTorch example:
import torch import torch.nn as nn import torch.optim as optim # Define two independent sub-models class FeatureModel(nn.Module): def __init__(self): super().__init__() self.linear = nn.Linear(1, 1) # Single input feature def forward(self, x): return self.linear(x) # Wrap them into a joint additive model class AdditiveJointModel(nn.Module): def __init__(self): super().__init__() self.model_x1 = FeatureModel() self.model_x2 = FeatureModel() def forward(self, x1, x2): pred1 = self.model_x1(x1) pred2 = self.model_x2(x2) return pred1 + pred2 # Prepare sample data (true relationship: y = 2*x1 + 3*x2 + noise) X = torch.randn(1000, 2) y = 2*X[:, 0].reshape(-1,1) + 3*X[:,1].reshape(-1,1) + torch.randn(1000,1)*0.1 # Initialize training components model = AdditiveJointModel() loss_fn = nn.MSELoss() optimizer = optim.SGD(model.parameters(), lr=0.01) # Training loop epochs = 1000 for epoch in range(epochs): optimizer.zero_grad() # Split features x1, x2 = X[:,0].reshape(-1,1), X[:,1].reshape(-1,1) # Compute combined predictions and loss total_pred = model(x1, x2) loss = loss_fn(total_pred, y) # Backpropagate loss and update all parameters loss.backward() optimizer.step() if epoch % 100 == 0: print(f"Epoch {epoch:4d} | Loss: {loss.item():.6f}") # Test prediction test_x = torch.tensor([[1.0, 2.0]]) test_pred = model(test_x[:,0].reshape(-1,1), test_x[:,1].reshape(-1,1)) print(f"\nTest Prediction: {test_pred.item():.4f} | True Value: {2*1 +3*2:.4f}")
Key Notes
- Never train separately: Training each model to predict
yalone will lead to suboptimal results, as they don’t account for the additive relationship. - Share the loss: The only way to ensure collaboration is to tie both models’ updates to the same loss function that evaluates their combined output.
- Choose the right tool: Use scikit-learn custom estimators for traditional ML models, and deep learning frameworks for neural networks (they handle gradient logic out of the box).
内容的提问来源于stack exchange,提问作者bernardo_soares

