You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于RNN时序反向传播中dh[t]与dh_prev求和的实现疑问

Why do we add dh[t] and dh_prev in RNN BPTT?

Great question—this line is the core of how backpropagation through time (BPTT) handles gradient flow in RNNs, so let’s unpack it with concrete intuition tied to the code you shared.

First, remember what each hidden state h[t] in an RNN does:

  • It directly contributes to the output at time step t (hence the gradient dh[t], which comes from the loss associated with that output).
  • It also acts as the "memory" that feeds into the next time step’s hidden state h[t+1]. When we reverse the computation for BPTT, the gradient from h[t+1] (which we’ve already calculated in the previous iteration of the reversed loop, stored in dh_prev) needs to flow back to h[t].

By the chain rule, the total gradient for h[t] is the sum of these two separate gradient sources:

  1. The gradient from the current time step’s output (dh[t])
  2. The gradient from the next time step’s hidden state (dh_prev)

That’s exactly what dh_current = dh[t] + dh_prev is doing: it combines both gradients to get the full, total gradient for h[t], which we then pass into rnn_step_backward to compute the gradients for the current step’s input, weights, and the gradient that will flow back to the previous time step (dh_prev gets updated here for the next iteration).

Let’s walk through a quick example with the loop:

  • When we start at the last time step t=T-1, dh_prev is 0 (since there’s no time step after T-1 to pass a gradient back from). So dh_current is just dh[T-1], the gradient from the final output.
  • After processing t=T-1, dh_prev is set to the gradient that flows back from h[T-1] to h[T-2].
  • When we move to t=T-2, we add dh[T-2] (gradient from this step’s output) to the dh_prev (gradient from h[T-1]), giving the total gradient for h[T-2]. This repeats until we reach the first time step.

Without this addition, we’d only account for one of the two gradient paths, leading to incorrect weight updates and a broken RNN training loop.

内容的提问来源于stack exchange,提问作者Ali Khalilli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:27:27