Keras中LSTM多对一架构的图像输入形状设置问题
Hey there! Let's break down how to set up the input_shape correctly for your LSTM multi-to-one architecture, since you're working with 19 sequential (128,128,3) image frames.
First, Let's Recall LSTM Input Structure
In Keras, LSTMs expect input in the format (batch_size, timesteps, data_dim)—but when defining input_shape in the layer, you omit the batch size (it's handled automatically during training).
You’re correct that each image frame has 128*128*3 = 49152 total pixels, but there are two common approaches to structure this input, depending on whether you use a standard LSTM or a more image-friendly ConvLSTM.
Approach 1: Standard LSTM (Flatten Each Frame)
If you stick with a regular LSTM, you need to flatten each 2D+channel image into a 1D vector first. This converts your input data from (batch_size, 19, 128, 128, 3) to (batch_size, 19, 49152).
Here’s how to adjust your code:
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import LSTM, Dense timesteps = 19 data_dim = 128 * 128 * 3 # 49152 total features per frame model = Sequential() # Set input_shape to (timesteps, data_dim) — no batch size needed model.add(LSTM(units=64, input_shape=(timesteps, data_dim))) # Adjust units based on your task # Add your output layer (e.g., Dense(10) for 10-class classification, Dense(1) for regression) model.add(Dense(units=1, activation='sigmoid')) # Example for binary classification model.compile(optimizer='adam', loss='binary_crossentropy')
Note: Flattening images loses spatial information (like edge or texture relationships), which isn’t ideal for image data. That’s why the next approach is often better.
Approach 2: ConvLSTM2D (Better for Image Sequences)
ConvLSTM2D is designed specifically for sequential image data—it preserves spatial features by applying convolutions across both space and time frames. This means you don’t need to flatten your images at all.
For a multi-to-one setup, set return_sequences=False (since we only want one final output, not outputs for each timestep):
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import ConvLSTM2D, Flatten, Dense model = Sequential() # input_shape here is (timesteps, height, width, channels) model.add(ConvLSTM2D( filters=32, kernel_size=(3, 3), input_shape=(19, 128, 128, 3), return_sequences=False # Critical for multi-to-one output )) # Flatten the ConvLSTM output to feed into a dense layer model.add(Flatten()) # Add your output layer model.add(Dense(units=1, activation='linear')) # Example for regression model.compile(optimizer='adam', loss='mse')
Key Takeaways
- Standard LSTM: Use
input_shape=(19, 49152)after flattening each frame. - ConvLSTM2D: Use
input_shape=(19, 128, 128, 3)to keep spatial information intact (recommended for image sequences).
内容的提问来源于stack exchange,提问作者Yaman K Singla

