如何将随机梯度下降(SGD)应用于正则化目标函数?
Great question! Let’s break this down clearly—regularization fits into SGD surprisingly neatly once you understand how the two pieces interact.
How SGD Adapts to Regularized Objectives
First, remember that vanilla SGD works by iteratively updating model parameters using the gradient of the loss function on a single sample (or small batch). Regularization adds an extra "penalty term" to the objective function to prevent overfitting—think of it as a way to nudge parameters toward simpler values (like zero, for most common regularization types).
The key adaptation here is super simple: instead of only computing the gradient of the sample loss, we also compute the gradient of the regularization penalty and add it to the total gradient before updating parameters. SGD’s core loop (sample → compute gradient → update parameters) stays intact—we just expand what "gradient" means to include the regularization term.
Applying SGD to Regularized Objective Functions
Let’s make this concrete with formulas and examples.
First, define the regularized objective function. Suppose our original unregularized loss is:
$$J(\theta) = \frac{1}{N}\sum_{i=1}^N Q_i(\theta)$$
where $Q_i(\theta)$ is the loss contribution from the $i$-th sample, and $\theta$ is our model’s parameter vector.
Adding regularization gives us:
$$J_{reg}(\theta) = \frac{1}{N}\sum_{i=1}^N Q_i(\theta) + \lambda R(\theta)$$
Here:
- $\lambda$ is the regularization strength (a hyperparameter that balances the sample loss and the penalty)
- $R(\theta)$ is the regularization term (common choices are L1 or L2)
Step-by-Step SGD Update for Regularized Objectives
For each iteration of SGD:
- Sample a mini-batch: Pick a small subset of samples $B$ (or a single sample, for vanilla SGD) from your training data.
- Compute sample loss gradient: Calculate the average gradient of $Q_i(\theta)$ over the mini-batch:
∇Q = (1/|B|) * sum(∇Q_i(θ) for i in B) - Compute regularization gradient: Calculate the gradient of the penalty term $R(\theta)$, scaled by $\lambda$:
∇R = λ * ∇R(θ) - Combine gradients: The total gradient for the regularized objective is the sum of the two:
∇J_reg = ∇Q + ∇R - Update parameters: Adjust $\theta$ using the total gradient and your learning rate $\eta$:
θ = θ - η * ∇J_reg
Examples for Common Regularization Types
Let’s apply this to the two most widely used regularization methods:
L2 Regularization (Weight Decay)
L2 regularization uses the squared magnitude of parameters as the penalty:
$$R(\theta) = \frac{1}{2}\sum_j \theta_j^2$$
The gradient of this term is straightforward: $\nabla R(\theta) = \theta$ (since the derivative of $\theta_j^2$ is $2\theta_j$, and the 1/2 cancels that out).
Plugging this into the SGD update gives:
$$θ_j = θ_j - η*(∇Q_j + λθ_j) = θ_j*(1 - ηλ) - η∇Q_j$$
You’ll notice this "decays" the parameter value by a factor of $(1 - ηλ)$ before applying the sample loss gradient—this is why L2 regularization is often called "weight decay."
L1 Regularization
L1 regularization uses the absolute magnitude of parameters as the penalty:
$$R(\theta) = \sum_j |θ_j|$$
The gradient here is the sign of each parameter: $\nabla R(\theta)_j = sign(θ_j)$ (for $θ_j ≠ 0$; when $θ_j = 0$, we use a subgradient, but in practice, $sign(0) = 0$ works fine).
The SGD update becomes:
$$θ_j = θ_j - η*(∇Q_j + λ*sign(θ_j))$$
L1 regularization tends to push parameters exactly to zero (creating sparse models), unlike L2 which just shrinks them toward zero.
Quick Implementation Tips
- Skip regularization for bias terms: Most practitioners don’t apply regularization to bias parameters (since they only shift predictions, not add complexity to the model). Make sure to exclude them when calculating $\nabla R$.
- Watch the scaling: Some frameworks include the 1/N factor in the regularization term, but it’s more common to leave that out and let $\lambda$ control the penalty strength. Always check your implementation’s conventions!
内容的提问来源于stack exchange,提问作者Jason Yu

