You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch LSTM中hidden(h_n)与output的区别及冗余性疑问

PyTorch LSTM: Output vs Hidden States (h_n, c_n)

Great question—this is a super common point of confusion when working with PyTorch's recurrent layers. Let's break this down clearly, starting with core differences, then addressing whether h_n is just a repeat of output's last step, and wrapping up with the unique uses of h_n.

1. Core Differences Between output and Hidden States

What is output?

output stores the output features of the last layer for every time step in your sequence. Its shape is (seq_len, batch, hidden_size * num_directions):

  • seq_len: Number of time steps in your input sequence
  • batch: Batch size
  • hidden_size * num_directions: For bidirectional models, this doubles the hidden size (we concatenate forward and backward step outputs)

In short: if your sequence has 10 time steps, output contains 10 separate h_t values—one for each moment t from 0 to 9—all coming from the final layer of the LSTM.

What are h_n and c_n?

These are the final hidden and cell states across all layers at the last time step (t = seq_len):

  • h_n: Shape (num_layers * num_directions, batch, hidden_size) → This holds the hidden state of every layer (and every direction, if bidirectional) at the end of the sequence.
  • c_n: Exclusive to LSTMs (not RNN/GRU), shape matches h_n → This is the cell state of every layer/direction at the final time step.

The key distinction here is "all layers": unlike output, which only gives you the final layer's outputs, h_n includes hidden states from every layer in your model.

2. Is h_n Just a Repeat of output's Last Time Step?

It depends entirely on your model configuration:

  • Single-layer, unidirectional LSTM: Yes! Here, h_n[0] (since num_layers * num_directions = 1) is identical to output[-1]. output[-1] is the final time step's hidden state from the only layer, which is exactly what h_n stores.
  • Multi-layer or bidirectional LSTM: No, not entirely.
    • For a 2-layer unidirectional LSTM, h_n has 2 entries: the final hidden state of layer 1, and the final hidden state of layer 2. output[-1] only matches the second entry (layer 2's final state)—the first entry (layer 1's final state) isn't present anywhere in output.
    • For a single-layer bidirectional LSTM, h_n has 2 entries: the final forward state, and the final backward state. output[-1] is the concatenation of these two states, so h_n[0] equals output[-1][:, :hidden_size] and h_n[1] equals output[-1][:, hidden_size:]—they're split components, not a direct repeat.

3. What Else is h_n Useful For?

Beyond being a shortcut to the final state in some cases, h_n has critical practical uses:

  • Initializing state for sequential batches: If you're processing a continuous stream (e.g., a long text split into chunks), you can pass h_n and c_n from the first batch as the initial hidden state for the next. This lets the model retain context across batches instead of resetting to zero every time.
  • Accessing intermediate layer states: For debugging, model analysis, or custom architectures (like hierarchical attention), you might need the final state of intermediate layers. h_n gives direct access to these, whereas output only provides the final layer's outputs.
  • Simplifying final-state tasks: For sequence classification (where you only care about overall sequence meaning), using h_n is cleaner than slicing output[-1]—especially for bidirectional/multi-layer models, where you might need to aggregate states from multiple layers/directions.

Quick Note on RNN/GRU

The logic applies similarly:

  • RNNs have output and h_n (no c_n, since they lack cell states).
  • GRUs also only have output and h_n—they merge cell and hidden states into one vector, so no separate c_n exists.

内容的提问来源于stack exchange,提问作者N. Virgo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:31:43