关于Keras中LSTM工作机制的技术问询及输出形状验证
Let's break down exactly why your model outputs a shape of (3,1) and walk through how the LSTM processes your input step-by-step.
First, Clarify the Key Parameters
Your input is a sequence of 3 words, each represented as a 4-dimensional vector — so the input shape (ignoring batch size) is (3,4). Here's what each part of your LSTM layer does:
LSTM(1): The1defines the dimension of the LSTM's hidden state output. Every time the LSTM processes a time step, it outputs a 1-dimensional vector representing its current state.return_sequences=True: This is the critical setting here. When enabled, the LSTM doesn't just return the hidden state from the final time step — it returns the hidden state generated at every single time step in the input sequence.
Why the Output Shape is (3,1)
Let's map this directly to your example sequence (how are you, 3 time steps):
- Time Step 1: The LSTM takes the first word's 4D vector, processes it, and outputs a 1D hidden state (
h1). - Time Step 2: The LSTM takes the second word's 4D vector plus the hidden state
h1from the previous step, processes them together, and outputs a new 1D hidden state (h2). - Time Step 3: The LSTM takes the third word's 4D vector plus
h2, processes them, and outputsh3.
Since return_sequences=True, the model stacks these three 1D states into a sequence: [h1, h2, h3], which gives you a shape of (3,1) (note: Keras actually shows (None, 3,1) in model.summary() — the None represents a variable batch size, but when feeding a single sample, it simplifies to (3,1)).
How This Fits Into a Translation Task
In translation scenarios, this setup makes sense if you're doing word-level sequence-to-sequence mapping (e.g., generating a translated word for every input word). If you instead wanted a single "summary" vector of the entire input sequence to feed into a decoder LSTM (the standard encoder-decoder translation setup), you'd set return_sequences=False (the default). That would give you only the final hidden state h3, with a shape of (1) (or (batch_size,1) for batches).
Quick Test: Adjusting return_sequences
If you modify your code to:
model.add(LSTM(1, input_shape=(3,4))) # return_sequences=False by default
Your output shape would change to (1) (or (batch_size,1)), since only the final time step's hidden state is returned.
内容的提问来源于stack exchange,提问作者Pusheen_the_dev

