You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

反向传播过程中是否始终使用Gradient Descent进行权重更新?使用Adam或RMSProp优化器时的权重更新机制疑问

Great questions! Let's break them down clearly:

1. 反向传播过程中是否始终使用Gradient Descent(梯度下降)进行权重更新?

Nope, that's a common mix-up between two distinct parts of model training! Let's set the record straight:

  • Backpropagation is only for gradient calculation: Its sole purpose is to compute the gradients of the loss function relative to each model parameter (weights, biases) using the chain rule. It doesn't handle parameter updates at all—it just passes the computed gradients to the optimizer.
  • Gradient Descent is one optimizer option: It's a basic method that uses those gradients to adjust parameters. But there are dozens of other optimizers (Adam, RMSprop, Adagrad, etc.) that are built on top of gradient descent ideas but add improvements like adaptive learning rates or momentum to make training faster and more stable.

In short: Backpropagation is the gradient-calculation step, and you can pair it with any optimizer—not just vanilla Gradient Descent.

2. 选用Adam/RMSprop优化器时,反向传播后用哪种方法更新权重?

When you run model.compile(optimizer = "adam") (or "rmsprop"), your model will use Adam (or RMSprop) directly for weight updates—not vanilla Gradient Descent.

Here's what that means in practice:

  • Vanilla Gradient Descent uses a simple update rule: weight = weight - learning_rate * gradient (fixed learning rate, only uses the current batch's gradients).
  • Adam takes this a step further with two key enhancements:
    1. Momentum: It maintains a moving average of past gradients to smooth out updates and avoid getting trapped in local minima.
    2. Adaptive learning rates: It adjusts the learning rate for each individual parameter based on how frequently its gradients change (so rarely-updated parameters get larger steps, and frequently-updated ones get smaller, more precise steps).
  • RMSprop works similarly but focuses only on adaptive learning rates (by normalizing gradients with a moving average of their squared values) without the momentum component.

Backpropagation still runs to compute the necessary gradients, but those gradients are fed into the Adam/RMSprop algorithm's update rules instead of the basic Gradient Descent formula.

内容的提问来源于stack exchange,提问作者pjrockzzz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 08:27:46