为何L2正则化会使代价函数更平滑?相关疑问解析
Great question—this is one of those concepts that clicks once you connect the math to real intuition, so let’s break it down simply:
Core Reason in Plain Language
L2 regularization smooths things out by penalizing large weights and adding a "dampening" term to the cost function that eliminates extreme, jagged fluctuations in both the function itself and how it changes (its gradients).
How It Works, Step by Step
1. It Tames Extreme Weights (and Their Wild Swings)
Without regularization, your model’s weights can balloon to huge values to overfit tiny noise in training data. When weights are massive, even tiny changes to them cause massive jumps in the cost function—like a steep rollercoaster where a tiny track shift sends you flying.
L2 fixes this by adding a term (λ/2) * sum(w_i²) to your original cost function J_original, where λ is the regularization strength and w_i are your model’s weights. This term grows quadratically with weight size, so it slaps a heavy penalty on any weight that tries to get too big. By forcing weights to stay small, the cost function becomes far less sensitive to tiny weight tweaks—no more sudden spikes or drops.
2. It Damps Gradients to Prevent Sharp Updates
When calculating gradients (to update weights during training), the L2 term adds λ*w to the original gradient. Here’s why that’s a game-changer:
- If a weight starts getting too large, the
λ*wterm pushes the gradient in the opposite direction, slowing its growth. - This means weight updates are never extreme, so the loss surface doesn’t develop those sharp, jagged peaks and valleys that come from unregulated weight growth.
3. Intuitive Analogy
Think of the unregularized loss surface as a rocky, uneven hiking trail. L2 regularization is like covering that trail with a thick layer of rubber matting—all the sharp rocks and steep drops get softened, making the whole path (surface) much smoother and easier to traverse (for your optimizer to find a stable minimum).
Bonus: The Technical Math Bit (For Curious Folks)
For those who want the deeper dive: The smoothness of a loss surface ties to its Hessian matrix (a matrix of second derivatives). L2 regularization adds λ*I (a diagonal matrix with λ on the diagonal) to the Hessian of the original cost function. This increases all the eigenvalues of the Hessian, which reduces the extreme curvature of the surface—so no more super steep slopes or sharp corners.
内容的提问来源于stack exchange,提问作者alwayscurious

