PyTorch LSTM中hidden(h_n)与output的区别及冗余性疑问
Great question—this is a super common point of confusion when working with PyTorch's recurrent layers. Let's break this down clearly, starting with core differences, then addressing whether h_n is just a repeat of output's last step, and wrapping up with the unique uses of h_n.
1. Core Differences Between output and Hidden States
What is output?
output stores the output features of the last layer for every time step in your sequence. Its shape is (seq_len, batch, hidden_size * num_directions):
seq_len: Number of time steps in your input sequencebatch: Batch sizehidden_size * num_directions: For bidirectional models, this doubles the hidden size (we concatenate forward and backward step outputs)
In short: if your sequence has 10 time steps, output contains 10 separate h_t values—one for each moment t from 0 to 9—all coming from the final layer of the LSTM.
What are h_n and c_n?
These are the final hidden and cell states across all layers at the last time step (t = seq_len):
h_n: Shape(num_layers * num_directions, batch, hidden_size)→ This holds the hidden state of every layer (and every direction, if bidirectional) at the end of the sequence.c_n: Exclusive to LSTMs (not RNN/GRU), shape matchesh_n→ This is the cell state of every layer/direction at the final time step.
The key distinction here is "all layers": unlike output, which only gives you the final layer's outputs, h_n includes hidden states from every layer in your model.
2. Is h_n Just a Repeat of output's Last Time Step?
It depends entirely on your model configuration:
- Single-layer, unidirectional LSTM: Yes! Here,
h_n[0](sincenum_layers * num_directions = 1) is identical tooutput[-1].output[-1]is the final time step's hidden state from the only layer, which is exactly whath_nstores. - Multi-layer or bidirectional LSTM: No, not entirely.
- For a 2-layer unidirectional LSTM,
h_nhas 2 entries: the final hidden state of layer 1, and the final hidden state of layer 2.output[-1]only matches the second entry (layer 2's final state)—the first entry (layer 1's final state) isn't present anywhere inoutput. - For a single-layer bidirectional LSTM,
h_nhas 2 entries: the final forward state, and the final backward state.output[-1]is the concatenation of these two states, soh_n[0]equalsoutput[-1][:, :hidden_size]andh_n[1]equalsoutput[-1][:, hidden_size:]—they're split components, not a direct repeat.
- For a 2-layer unidirectional LSTM,
3. What Else is h_n Useful For?
Beyond being a shortcut to the final state in some cases, h_n has critical practical uses:
- Initializing state for sequential batches: If you're processing a continuous stream (e.g., a long text split into chunks), you can pass
h_nandc_nfrom the first batch as the initial hidden state for the next. This lets the model retain context across batches instead of resetting to zero every time. - Accessing intermediate layer states: For debugging, model analysis, or custom architectures (like hierarchical attention), you might need the final state of intermediate layers.
h_ngives direct access to these, whereasoutputonly provides the final layer's outputs. - Simplifying final-state tasks: For sequence classification (where you only care about overall sequence meaning), using
h_nis cleaner than slicingoutput[-1]—especially for bidirectional/multi-layer models, where you might need to aggregate states from multiple layers/directions.
Quick Note on RNN/GRU
The logic applies similarly:
- RNNs have
outputandh_n(noc_n, since they lack cell states). - GRUs also only have
outputandh_n—they merge cell and hidden states into one vector, so no separatec_nexists.
内容的提问来源于stack exchange,提问作者N. Virgo

