关于RNN时序反向传播中dh[t]与dh_prev求和的实现疑问
dh[t] and dh_prev in RNN BPTT? Great question—this line is the core of how backpropagation through time (BPTT) handles gradient flow in RNNs, so let’s unpack it with concrete intuition tied to the code you shared.
First, remember what each hidden state h[t] in an RNN does:
- It directly contributes to the output at time step
t(hence the gradientdh[t], which comes from the loss associated with that output). - It also acts as the "memory" that feeds into the next time step’s hidden state
h[t+1]. When we reverse the computation for BPTT, the gradient fromh[t+1](which we’ve already calculated in the previous iteration of the reversed loop, stored indh_prev) needs to flow back toh[t].
By the chain rule, the total gradient for h[t] is the sum of these two separate gradient sources:
- The gradient from the current time step’s output (
dh[t]) - The gradient from the next time step’s hidden state (
dh_prev)
That’s exactly what dh_current = dh[t] + dh_prev is doing: it combines both gradients to get the full, total gradient for h[t], which we then pass into rnn_step_backward to compute the gradients for the current step’s input, weights, and the gradient that will flow back to the previous time step (dh_prev gets updated here for the next iteration).
Let’s walk through a quick example with the loop:
- When we start at the last time step
t=T-1,dh_previs 0 (since there’s no time step afterT-1to pass a gradient back from). Sodh_currentis justdh[T-1], the gradient from the final output. - After processing
t=T-1,dh_previs set to the gradient that flows back fromh[T-1]toh[T-2]. - When we move to
t=T-2, we adddh[T-2](gradient from this step’s output) to thedh_prev(gradient fromh[T-1]), giving the total gradient forh[T-2]. This repeats until we reach the first time step.
Without this addition, we’d only account for one of the two gradient paths, leading to incorrect weight updates and a broken RNN training loop.
内容的提问来源于stack exchange,提问作者Ali Khalilli

