You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何损失在多epoch恒定时需降低学习率?直观解释与局部极小值疑问

Why Lower Learning Rate When Loss Stays Constant Across Epochs?

Intuitive Explanation

Think of training a neural network like hunting for the lowest valley in a rugged mountain range (the valley represents the minimal loss we’re chasing):

  • Your learning rate is the length of each step you take. If your steps are too big, when you get close to the valley, you’ll likely overshoot it—stepping to the other side, then stepping back the next epoch. You’re stuck oscillating right around the valley, never actually dropping into it, so your loss stays flat.
  • When loss holds steady across multiple epochs, it’s almost always a sign of this "oscillatory stall": the current learning rate is too large, so each parameter update just bounces the model between points with the same loss level. Shrinking the step size (lowering the learning rate) lets the model "creep" into the valley, capturing smaller drops in loss that the big steps were skipping over.

Another way to look at it: a flat loss means your current parameter updates aren’t moving the needle anymore. Either the updates are too large to make meaningful progress, or they’re canceling each other out. Lowering the learning rate switches the model from "rough exploration" to "fine-grained search," letting it dig into smaller improvements.

Does Adjusting Learning Rate Lead to Local Minima or Plateaus?

Short answer: Smart, targeted adjustments won’t cause local minima or plateaus—in fact, they often help break fake plateaus.

Let’s unpack this:

  • Local minima: Lowering the learning rate doesn’t create local minima. If your model is already stuck in a local minimum, a tiny learning rate might make it harder to escape—but that’s a problem you had before adjusting the rate. If the flat loss was from oscillation (a fake plateau), reducing the learning rate will let the model settle into a true minimum (whether local or global) that it couldn’t reach with big steps.
  • Plateaus: There are two types to distinguish:
    • Fake plateaus: Caused by large learning rate oscillations. The loss looks flat, but the model still has room to improve. Lowering the learning rate breaks this cycle immediately, letting loss start dropping again.
    • True plateaus: These are flat regions in the loss function itself, where no parameter update will move the needle. This isn’t caused by adjusting the learning rate—it’s a property of the model architecture, dataset, or task. If you hit a true plateau, lowering the learning rate won’t fix it, but it also didn’t cause it. You’ll need other fixes like changing the model, adding data augmentation, or tweaking regularization.

One caveat: Don’t drop the learning rate too much too fast. If your step size becomes tiny, the model might get stuck in a suboptimal spot, unable to even move to a slightly better nearby point. Gradual decays (like cutting the rate by 10% every 5 epochs, or using cosine annealing) are usually safer than drastic one-time drops.


内容的提问来源于stack exchange,提问作者MenorcanOrange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:46:45