You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练集与验证集学习曲线重合?scikit-learn训练问题咨询

Troubleshooting MLPRegressor Learning Curves That Don't Separate

Hey Andy, let's dive into why your MLPRegressor's training and validation learning curves aren't showing the distinct separation you'd expect, and how to diagnose what's going on.

Common Reasons for Overlapping Learning Curves

First, let's break down the most likely causes:

  • Underfitting (Model Too Simple)
    If your MLP is too small (e.g., too few hidden layers/neurons) or heavily regularized (high alpha), it might not have enough capacity to capture complex patterns in your data. In this case, both training and validation errors stay high and close together because the model can't even fit the training data well.

  • Data Is "Easy" to Model
    Your dataset might have a very clear, linear relationship between features and target, or extremely low noise. In this scenario, even a simple model (including your tuned MLP) can generalize nearly perfectly, so training and validation errors overlap because there's little room for overfitting.

  • GridSearchCV Tuned for Underfitting
    It's possible your parameter search landed on settings that prioritize generalization at the cost of model capacity—like a very high alpha value or linear activation (activation='identity'). This would push the model into an underfitting regime.

  • MAPE Scoring Quirks
    Mean Absolute Percentage Error (MAPE) behaves differently from metrics like MSE: it's sensitive to small target values, and if your target range is narrow, small error differences might not show up visually as distinct curves.

Step-by-Step Diagnosis & Fixes

1. Check Your GridSearch Results

First, confirm what parameters GridSearchCV selected—this will tell you if the model is intentionally constrained:

print("Best MLP Parameters:", grid_search.best_params_)
print("Best Cross-Validation MAPE:", -grid_search.best_score_)  # Reverse if you used greater_is_better=False

If you see high alpha, tiny hidden layers, or linear activation, that's a sign the model is underfitting.

2. Expand Your Parameter Grid

Try increasing model capacity and adjusting regularization to give the MLP more room to learn complex patterns:

param_grid = {
    'hidden_layer_sizes': [(50,), (100,), (100, 50), (200, 100)],  # Add deeper/wider layers
    'alpha': [0.0001, 0.001, 0.01, 0.1],  # Broaden regularization range
    'activation': ['relu', 'tanh'],  # Try non-linear activations
    'solver': ['adam', 'sgd'],  # Test different optimizers
    'max_iter': [500, 1000]  # Ensure model has enough epochs to converge
}

3. Verify Your Learning Curve Code

Double-check that your MAPE scorer and learning curve implementation are correct (it's easy to mix up sign conventions):

import numpy as np
import matplotlib.pyplot as plt
from sklearn.metrics import make_scorer
from sklearn.model_selection import learning_curve, ShuffleSplit

# Custom MAPE scorer (note the sign reversal for GridSearch/learning_curve)
def mape_score(y_true, y_pred):
    return np.mean(np.abs((y_true - y_pred) / y_true)) * 100

mape_scorer = make_scorer(mape_score, greater_is_better=False)

# Generate learning curves with your ShuffleSplit setup
cv = ShuffleSplit(n_splits=10, test_size=0.2, random_state=42)
train_sizes, train_scores, val_scores = learning_curve(
    grid_search.best_estimator_,
    X, y,
    cv=cv,
    scoring=mape_scorer,
    train_sizes=np.linspace(0.1, 1.0, 10)
)

# Convert scores back to positive MAPE for plotting
train_mape = -train_scores.mean(axis=1)
val_mape = -val_scores.mean(axis=1)
train_std = train_scores.std(axis=1)
val_std = val_scores.std(axis=1)

# Plot the curves
plt.figure(figsize=(10, 6))
plt.title("MLPRegressor Learning Curve (MAPE)")
plt.xlabel("Number of Training Samples")
plt.ylabel("MAPE (%)")
plt.grid(True)

plt.fill_between(train_sizes, train_mape - train_std, train_mape + train_std, alpha=0.1, color="red")
plt.fill_between(train_sizes, val_mape - val_std, val_mape + val_std, alpha=0.1, color="green")
plt.plot(train_sizes, train_mape, 'o-', color="red", label="Training MAPE")
plt.plot(train_sizes, val_mape, 'o-', color="green", label="Validation MAPE")

plt.legend(loc="best")
plt.show()

4. Test with a Different Metric

To rule out MAPE as the culprit, generate learning curves using MSE or MAE instead. If the curves separate with these metrics, the issue is just how MAPE visualizes your data's error.

Final Thought

If after all these steps your curves still overlap, don't panic—it might mean your model is actually performing fantastically! If both training and validation MAPE are low and close, your MLP is generalizing perfectly to unseen data. The "separated curves" you see in examples are usually cases where overfitting is happening, which isn't always the norm.

内容的提问来源于stack exchange,提问作者Andy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:56:27