You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

深度学习优化:批量、小批量及随机梯度下降技术问询

深度学习优化入门:以线性回归为例

Hey everyone, let's dive into optimization in deep learning using linear regression as a concrete example—it's a great starting point since it's intuitive and lays the groundwork for more complex methods.

先明确核心符号

Before we jump into the algorithm, let's define a few key terms to avoid confusion:

  • $m$:训练集$(X,y)$的总数据点数
  • $n$:迭代次数(也就是我们让优化算法跑多少轮)

批量梯度下降(GD):从原理到代码

Batch Gradient Descent (GD) is the foundational optimization algorithm here. The core idea is straightforward: we use the entire training dataset to compute the gradient in every iteration.

为什么要用向量化操作?

If you've ever written a loop to compute gradients manually, you know how slow it can get—especially with large datasets. That's where vectorization comes in! Using NumPy's np.dot(X.T, y_hat - y) lets us leverage optimized C under the hood, which is way faster than Python loops. It's a must-have trick for efficient ML code.

完整代码逻辑

Here's the step-by-step flow for batch GD with linear regression:

  1. First, compute the predicted values $y_hat$ for all training samples using the current parameters $\theta$: $y_hat = X \cdot \theta$ (assuming $X$ includes a bias column of 1s)
  2. Calculate the gradient across the entire dataset with np.dot(X.T, y_hat - y)—this gives us the sum of the gradients from every sample
  3. Normalize the gradient by dividing by $m$ (the number of samples) to get the average gradient
  4. Update the parameters $\theta$ using the learning rate $\alpha$: $\theta = \theta - \alpha \cdot \text{(average gradient)}$
  5. Repeat steps 1-4 for $n$ iterations, or until the gradient becomes small enough (we say the model has "converged")

代码示例

Here's a simple, vectorized implementation in Python:

import numpy as np

def batch_gradient_descent(X, y, theta_init, alpha, n_iterations):
    m = len(y)
    theta = theta_init.copy()
    for _ in range(n_iterations):
        # Compute predictions for all samples
        y_hat = np.dot(X, theta)
        # Calculate gradient using vectorization
        gradient = np.dot(X.T, y_hat - y) / m
        # Update parameters
        theta -= alpha * gradient
    return theta

关键优缺点

Let's wrap up with a quick pros/cons check for batch GD:

  • Pros: The gradient estimate is super accurate because we use all data, so the loss function decreases smoothly. It's also easy to implement and understand.
  • Cons: When $m$ is large (like millions of samples), computing the gradient across all data every iteration is slow and memory-heavy. This makes batch GD impractical for big datasets.

内容的提问来源于stack exchange,提问作者KevinKim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:32:01