深度学习优化:批量、小批量及随机梯度下降技术问询
Hey everyone, let's dive into optimization in deep learning using linear regression as a concrete example—it's a great starting point since it's intuitive and lays the groundwork for more complex methods.
先明确核心符号
Before we jump into the algorithm, let's define a few key terms to avoid confusion:
- $m$:训练集$(X,y)$的总数据点数
- $n$:迭代次数(也就是我们让优化算法跑多少轮)
批量梯度下降(GD):从原理到代码
Batch Gradient Descent (GD) is the foundational optimization algorithm here. The core idea is straightforward: we use the entire training dataset to compute the gradient in every iteration.
为什么要用向量化操作?
If you've ever written a loop to compute gradients manually, you know how slow it can get—especially with large datasets. That's where vectorization comes in! Using NumPy's np.dot(X.T, y_hat - y) lets us leverage optimized C under the hood, which is way faster than Python loops. It's a must-have trick for efficient ML code.
完整代码逻辑
Here's the step-by-step flow for batch GD with linear regression:
- First, compute the predicted values $y_hat$ for all training samples using the current parameters $\theta$: $y_hat = X \cdot \theta$ (assuming $X$ includes a bias column of 1s)
- Calculate the gradient across the entire dataset with
np.dot(X.T, y_hat - y)—this gives us the sum of the gradients from every sample - Normalize the gradient by dividing by $m$ (the number of samples) to get the average gradient
- Update the parameters $\theta$ using the learning rate $\alpha$: $\theta = \theta - \alpha \cdot \text{(average gradient)}$
- Repeat steps 1-4 for $n$ iterations, or until the gradient becomes small enough (we say the model has "converged")
代码示例
Here's a simple, vectorized implementation in Python:
import numpy as np def batch_gradient_descent(X, y, theta_init, alpha, n_iterations): m = len(y) theta = theta_init.copy() for _ in range(n_iterations): # Compute predictions for all samples y_hat = np.dot(X, theta) # Calculate gradient using vectorization gradient = np.dot(X.T, y_hat - y) / m # Update parameters theta -= alpha * gradient return theta
关键优缺点
Let's wrap up with a quick pros/cons check for batch GD:
- Pros: The gradient estimate is super accurate because we use all data, so the loss function decreases smoothly. It's also easy to implement and understand.
- Cons: When $m$ is large (like millions of samples), computing the gradient across all data every iteration is slow and memory-heavy. This makes batch GD impractical for big datasets.
内容的提问来源于stack exchange,提问作者KevinKim

