You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

梯度下降中学习率(步长)与动量参数的区别咨询

Learning Rate vs. Momentum in Gradient Descent

Great question—these two parameters are easy to mix up, but they serve totally different roles in making gradient descent work well. Let’s break them down with simple analogies and clear explanations:

Core Role of Learning Rate

The learning rate (usually denoted η or α) is all about how big a step you take each iteration. Think of it like hiking down a mountain to find the lowest valley:

  • If your learning rate is too large, you’ll take giant strides that might overshoot the valley—you could end up bouncing back and forth across the lowest point, or even climbing back up the other side (diverging instead of converging).
  • If it’s too small, you’ll creep along slowly, taking forever to reach the bottom, or getting stuck in a shallow dip before you find the real minimum.

Mathematically, it directly scales the current gradient of your loss function:
θ_new = θ_old - η * ∇L(θ_old)
It’s a scalar that applies the same scaling to all parameter updates (unless you’re using adaptive methods like Adam, but that’s a separate topic).

Core Role of Momentum

Momentum (often denoted β) is all about using past gradient information to smooth out your path and build "inertia". Going back to the mountain analogy:

  • Without momentum, each step only cares about the slope under your feet right now. If the path is rocky (noisy gradients or a bumpy loss surface), you’ll stumble back and forth, wasting steps.
  • With momentum, you’re like a rolling boulder—you carry forward the direction you’ve been moving in. If you’ve been heading downhill for a few steps, that momentum keeps you going in that direction, even if a small rocky patch tries to nudge you off course.

Mathematically, it maintains a running average of past gradients to compute the update direction:
v_t = β * v_{t-1} + (1 - β) * ∇L(θ_t)
θ_new = θ_old - η * v_t
Here, β controls how much weight you give to past updates (usually around 0.9). It helps you:

  • Speed up convergence by keeping you moving in a consistent direction
  • Reduce oscillations on noisy loss surfaces
  • Escape shallow local minima more easily

Key Distinction at a Glance

  • Learning Rate: Determines the magnitude of each step (how far you move).
  • Momentum: Determines the direction stability of each step (how much you rely on past movement to stay on track).

You can think of them as a team: the learning rate sets how hard you push, and momentum makes sure that push is consistent instead of erratic.

内容的提问来源于stack exchange,提问作者Katherine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:53:49