You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在scikit-learn中绘制代价与迭代次数图?梯度下降模型是否支持获取对应数据?

Great question! Let’s break this down step by step—since debugging gradient descent via cost function curves is exactly the kind of practical tuning Andrew Ng emphasizes in his Coursera ML course, and it’s totally doable in scikit-learn.

How to Plot Cost Function Curves for Gradient Descent Models in scikit-learn

First: Which scikit-learn Models Use Gradient Descent?

scikit-learn’s primary gradient descent implementations are found in the SGDRegressor (for regression tasks) and SGDClassifier (for classification) classes. These support batch, mini-batch, and stochastic gradient descent via simple parameter tweaks. Additionally, models like LogisticRegression with solvers 'sag' or 'saga', and Perceptron, also use gradient-based methods, but the SGD family is the most straightforward for accessing cost function data.

Getting Cost vs. Iteration Data

For SGDRegressor/SGDClassifier (The Easy Path)

The SGD models have a built-in loss_curve_ attribute that gets populated after training. This list contains the average cost (J(θ)) for each training epoch (full pass over the entire dataset).

A few key notes to make this work well:

  • Set max_iter to a sufficiently high value (like 1000) and tol to a small threshold (or None if you want to force all iterations) to avoid early stopping before you collect enough data points for your curve.
  • The loss_curve_ tracks average loss per epoch, not per individual gradient step (unless you’re using stochastic gradient descent with a batch size of 1, which is rare for this debugging purpose).

For Custom/Finer-Grained Tracking

If you need to track loss per mini-batch instead of per epoch, you’ll need to build a manual training loop using partial_fit():

  • Split your data into mini-batches of your desired size.
  • Initialize the model with warm_start=True so it retains learned weights between calls to partial_fit().
  • For each mini-batch, call partial_fit(), then calculate the current cost using the model’s loss function, and append it to a list along with the iteration count.

Plotting the Cost Curve (Full Code Example)

Let’s use a regression task with the diabetes dataset to demonstrate (standardization is critical for gradient descent to work properly—don’t skip this step!):

import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import SGDRegressor
from sklearn.datasets import load_diabetes
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split

# Load and split data
X, y = load_diabetes(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Standardize features (gradient descent performs poorly without scaled features!)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

# Initialize SGD Regressor
# Set tol=None to force training for all max_iter epochs (great for debugging)
sgd_reg = SGDRegressor(
    loss='squared_error',  # Matches the J(θ) cost function from linear regression
    max_iter=1000,
    tol=None,
    eta0=0.01,  # This is your learning rate α
    learning_rate='constant',  # Keep α fixed to test its impact
    random_state=42
)

# Train the model
sgd_reg.fit(X_train_scaled, y_train)

# Extract cost and iteration data
cost_history = sgd_reg.loss_curve_
iterations = np.arange(1, len(cost_history) + 1)

# Plot the curve
plt.figure(figsize=(10, 6))
plt.plot(iterations, cost_history, 'b-', linewidth=2)
plt.xlabel('Number of Iterations (Epochs)')
plt.ylabel('Cost Function J(θ)')
plt.title('Cost Function vs. Iteration for Gradient Descent')
plt.grid(True, alpha=0.3)
plt.show()

What If J(θ) Is Increasing?

Just like Andrew Ng teaches, if your cost curve is rising instead of decreasing, the most likely fix is to reduce your learning rate α (the eta0 parameter in SGD models). Try lowering eta0 to 0.001, 0.0001, etc., and re-run the training. You can also switch to a dynamic learning rate (like learning_rate='optimal') which automatically adjusts α over time to prevent divergence.

For Other Gradient-Based Models (e.g., LogisticRegression with SAG/SAGA)

Models like LogisticRegression using solver='sag' don’t have a built-in loss_curve_ attribute. To track cost here, you’ll need to:

  1. Initialize the model with warm_start=True and max_iter=1 (so it runs one epoch per fit call).
  2. Loop over your desired number of iterations, calling fit() each time.
  3. After each fit, calculate the current loss manually using the model’s predictions and the appropriate loss function (e.g., cross-entropy for classification).
  4. Append each loss value to your history list for plotting later.

内容的提问来源于stack exchange,提问作者Chris Snow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:47:15