LSTM网络结构图解及时间序列预测模型相关疑问咨询
Great question—let's break this down clearly, since LSTM's structure can feel counterintuitive when coming from standard feedforward neural networks.
1. Clarifying LSTM Units vs. Time Steps
First, let's correct a common misunderstanding about your network:
You asked: "是否包含20个LSTM细胞且每个细胞配备300个单元?"
No, that's not how it works. YourLSTM(300)layer is a single layer containing 300 hidden units (think of these as 300 parallel LSTM "cells"), and these 300 units process all 20 of your input time steps sequentially:
- Your input shape
(20, 1)means each training sample is a sequence of 20 time steps, with 1 feature per time step (e.g.,x_t-20tox_t). - For each of the 20 time steps, the 300 LSTM units update their internal cell state and hidden state using the current time step's input and the previous step's state.
- Since you didn't set
return_sequences=True(it defaults toFalse), this LSTM layer only outputs the hidden state from the final 20th time step—a 300-dimensional vector—rather than outputting a state for every time step.
2. How Your 20 Future Steps Are Generated
Your network uses a direct multi-step forecasting approach, which is straightforward for your task:
model = Sequential() model.add(LSTM(units=300, activation=activation, input_shape=(20, 1))) # Default: return_sequences=False model.add(Dense(20)) # Outputs 20 values for x_t+1 to x_t+20
Here's the step-by-step flow:
- Your input sequence (20 time steps) feeds into the LSTM layer. The layer processes each time step in order, updating the 300 units' states along the way.
- After processing the final time step (
x_t), the LSTM outputs its 300-dimensional hidden state. - This single vector is passed to the
Dense(20)layer, which maps it directly to 20 output values—each corresponding to one of your future time steps (x_t+1throughx_t+20).
So yes, all 20 future predictions are generated from the final LSTM hidden state.
3. Network Structure Diagram (Text-Based)
To visualize the flow, here's a simplified breakdown:
Input Sequence: [x_t-20, x_t-19, ..., x_t] (20 time steps, 1 feature each) | ▼ LSTM Layer (300 hidden units, return_sequences=False) | ▼ Only outputs the final time step's hidden state Final Hidden State: [h_t] (300-dimensional vector) | ▼ Dense Layer (20 output units) | ▼ Predicted Sequence: [x_t+1, x_t+2, ..., x_t+20] (20 future time steps)
If we zoom into the LSTM's internal time step processing (simplified):
Time Step 1: x_t-20 → Update 300 units' states → Generate h_t-20 (not output) Time Step 2: x_t-19 + h_t-20 → Update states → Generate h_t-19 (not output) ... Time Step 20: x_t + h_t-1 → Update states → Generate h_t (sent to Dense layer)
Quick Side Note: Alternative Architectures
If you ever wanted the LSTM to output a state for every time step (e.g., to stack another LSTM layer), you'd set return_sequences=True. But for your direct multi-step forecasting task, using the final state to map to all 20 future steps is a perfectly valid, simple, and efficient choice.
内容的提问来源于stack exchange,提问作者Gunay Abdullayeva

