You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何联合优化两个scikit-learn模型的特征拆分加性预测拟合?

Great question! The key here is that you don’t want to train the two models separately—because that would optimize each model to predict y on its own, not their sum. Instead, you need to jointly optimize both models by tying their updates to a single loss function that measures how well their combined predictions match y.

Let’s break down how to do this with practical implementations:


Core Idea

Instead of minimizing loss for each model individually (e.g., MSE(f1(x1), y) and MSE(f2(x2), y)), we minimize a single loss that evaluates the sum of their predictions against y:

Loss = MSE(f1(x1) + f2(x2), y)

Both models’ parameters are updated in lockstep based on this shared loss, ensuring they collaborate to fit the additive relationship.


Method 1: Custom Scikit-Learn Estimator (for traditional ML models)

Scikit-learn doesn’t have built-in support for this kind of joint training, but you can create a custom estimator to wrap two models and train them together. Here’s an example using SGDRegressor (works with any model that exposes its parameters):

from sklearn.base import BaseEstimator, RegressorMixin
from sklearn.linear_model import SGDRegressor
import numpy as np

class AdditiveJointModel(BaseEstimator, RegressorMixin):
    def __init__(self, model1=None, model2=None):
        # Use SGDRegressor by default, or let users pass custom models
        self.model1 = model1 if model1 else SGDRegressor(eta0=0.01)
        self.model2 = model2 if model2 else SGDRegressor(eta0=0.01)
        
    def fit(self, X, y, epochs=100, batch_size=32):
        # Split features into x1 and x2
        X1 = X[:, 0].reshape(-1, 1)
        X2 = X[:, 1].reshape(-1, 1)
        n_samples = X.shape[0]
        
        for _ in range(epochs):
            # Shuffle data for stochastic training
            indices = np.random.permutation(n_samples)
            X1_shuffled, X2_shuffled, y_shuffled = X1[indices], X2[indices], y[indices]
            
            # Train in mini-batches
            for i in range(0, n_samples, batch_size):
                end = min(i + batch_size, n_samples)
                x1_batch, x2_batch, y_batch = X1_shuffled[i:end], X2_shuffled[i:end], y_shuffled[i:end]
                
                # Compute combined predictions
                pred1 = self.model1.predict(x1_batch)
                pred2 = self.model2.predict(x2_batch)
                total_pred = pred1 + pred2
                
                # Calculate gradient for MSE loss (dLoss/dPred = 2*(total_pred - y_batch))
                error = total_pred - y_batch
                grad1 = 2 * error.reshape(-1, 1) * x1_batch
                grad2 = 2 * error.reshape(-1, 1) * x2_batch
                
                # Update model parameters manually (matches SGD logic)
                self.model1.coef_ -= self.model1.eta0 * np.mean(grad1, axis=0)
                self.model1.intercept_ -= self.model1.eta0 * np.mean(2 * error)
                
                self.model2.coef_ -= self.model2.eta0 * np.mean(grad2, axis=0)
                self.model2.intercept_ -= self.model2.eta0 * np.mean(2 * error)
        
        return self
    
    def predict(self, X):
        X1 = X[:, 0].reshape(-1, 1)
        X2 = X[:, 1].reshape(-1, 1)
        return self.model1.predict(X1) + self.model2.predict(X2)

Method 2: Deep Learning Framework (PyTorch/TensorFlow, more flexible)

For neural networks or complex models, using a deep learning framework simplifies joint optimization—since it handles gradient computation and parameter updates automatically. Here’s a PyTorch example:

import torch
import torch.nn as nn
import torch.optim as optim

# Define two independent sub-models
class FeatureModel(nn.Module):
    def __init__(self):
        super().__init__()
        self.linear = nn.Linear(1, 1)  # Single input feature
        
    def forward(self, x):
        return self.linear(x)

# Wrap them into a joint additive model
class AdditiveJointModel(nn.Module):
    def __init__(self):
        super().__init__()
        self.model_x1 = FeatureModel()
        self.model_x2 = FeatureModel()
        
    def forward(self, x1, x2):
        pred1 = self.model_x1(x1)
        pred2 = self.model_x2(x2)
        return pred1 + pred2

# Prepare sample data (true relationship: y = 2*x1 + 3*x2 + noise)
X = torch.randn(1000, 2)
y = 2*X[:, 0].reshape(-1,1) + 3*X[:,1].reshape(-1,1) + torch.randn(1000,1)*0.1

# Initialize training components
model = AdditiveJointModel()
loss_fn = nn.MSELoss()
optimizer = optim.SGD(model.parameters(), lr=0.01)

# Training loop
epochs = 1000
for epoch in range(epochs):
    optimizer.zero_grad()
    
    # Split features
    x1, x2 = X[:,0].reshape(-1,1), X[:,1].reshape(-1,1)
    
    # Compute combined predictions and loss
    total_pred = model(x1, x2)
    loss = loss_fn(total_pred, y)
    
    # Backpropagate loss and update all parameters
    loss.backward()
    optimizer.step()
    
    if epoch % 100 == 0:
        print(f"Epoch {epoch:4d} | Loss: {loss.item():.6f}")

# Test prediction
test_x = torch.tensor([[1.0, 2.0]])
test_pred = model(test_x[:,0].reshape(-1,1), test_x[:,1].reshape(-1,1))
print(f"\nTest Prediction: {test_pred.item():.4f} | True Value: {2*1 +3*2:.4f}")

Key Notes

  • Never train separately: Training each model to predict y alone will lead to suboptimal results, as they don’t account for the additive relationship.
  • Share the loss: The only way to ensure collaboration is to tie both models’ updates to the same loss function that evaluates their combined output.
  • Choose the right tool: Use scikit-learn custom estimators for traditional ML models, and deep learning frameworks for neural networks (they handle gradient logic out of the box).

内容的提问来源于stack exchange,提问作者bernardo_soares

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:54:43