You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:Sklearn中MLP早停与正则化的实现方案

Got it, let's fix this overfitting problem with Scikit-learn's MLP—this is such a common pain point, and you’re exactly right that regularization and early stopping are the key solutions. Let’s walk through a concrete, actionable implementation step by step:

1. First: Don’t Forget Data Preprocessing (Critical for MLP!)

Before jumping into regularization, make sure you’re scaling your data. MLPs are extremely sensitive to feature scales, and unnormalized data can lead to unstable training, poor convergence, or even worse overfitting. Here’s how to add standardization:

from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split

# Split your data first (replace with your actual dataset)
X_train, X_test, y_train, y_test = train_test_split(your_X, your_y, test_size=0.2, random_state=42)

# Scale features to mean=0, std=1
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
2. Implement L2 Regularization (Weight Decay)

Scikit-learn’s MLP uses the alpha parameter to control L2 regularization—this adds a penalty term to the loss function that discourages large weights (the main driver of overfitting). Start with a small value and adjust based on results:

from sklearn.neural_network import MLPClassifier  # Use MLPRegressor for regression tasks

# Initialize MLP with L2 regularization
mlp_reg = MLPClassifier(
    hidden_layer_sizes=(100, 50),  # Adjust based on your task (fewer neurons = less overfitting)
    alpha=0.001,  # Start here; try 0.0001, 0.01, or 0.1 if needed
    max_iter=500,
    random_state=42
)

# Train on scaled data
mlp_reg.fit(X_train_scaled, y_train)

# Check scores
print(f"Training set score: {mlp_reg.score(X_train_scaled, y_train):.4f}")
print(f"Test set score: {mlp_reg.score(X_test_scaled, y_test):.4f}")
  • Pro tip: If your test score is still low, increase alpha to strengthen regularization. If your training score drops too much (e.g., below 0.6), decrease alpha to avoid underfitting.
3. Add Early Stopping

Early stopping halts training when the model’s performance on a validation set stops improving—this prevents the model from memorizing noise in the training data. Combine it with regularization for best results:

# Initialize MLP with both regularization and early stopping
mlp_best = MLPClassifier(
    hidden_layer_sizes=(100, 50),
    alpha=0.001,
    early_stopping=True,  # Enable early stopping
    validation_fraction=0.1,  # Use 10% of training data as validation set
    n_iter_no_change=10,  # Stop after 10 epochs with no improvement
    tol=1e-4,  # Minimum improvement required to continue
    max_iter=500,  # Set a high max so early stopping can trigger first
    random_state=42
)

mlp_best.fit(X_train_scaled, y_train)

# Check results
print(f"Training set score: {mlp_best.score(X_train_scaled, y_train):.4f}")
print(f"Test set score: {mlp_best.score(X_test_scaled, y_test):.4f}")
print(f"Actual epochs trained: {mlp_best.n_iter_}")  # See how early it stopped
  • Adjust validation_fraction to 0.2 if you have a large training dataset—this gives a more reliable validation signal.
  • Tweak n_iter_no_change to 5 or 15 depending on how quickly you want to stop training.
4. Bonus: Hyperparameter Tuning

If you want to find the optimal alpha and hidden layer sizes, use GridSearchCV to automate testing:

from sklearn.model_selection import GridSearchCV

param_grid = {
    'alpha': [0.0001, 0.001, 0.01, 0.1],
    'hidden_layer_sizes': [(50,), (100,), (100, 50)]
}

grid_search = GridSearchCV(MLPClassifier(early_stopping=True, max_iter=500, random_state=42),
                           param_grid, cv=5, scoring='accuracy')

grid_search.fit(X_train_scaled, y_train)

print(f"Best parameters: {grid_search.best_params_}")
print(f"Best cross-validation score: {grid_search.best_score_:.4f}")
print(f"Test set score with best model: {grid_search.score(X_test_scaled, y_test):.4f}")

Remember, the goal is to balance training set performance with test set performance—you don’t need a perfect training score, just one that generalizes well to unseen data.

内容的提问来源于stack exchange,提问作者100kasimpasali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:50:13