You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras中LSTM多对一架构的图像输入形状设置问题

How to Set Input Shape for LSTM Multi-to-One with Image Frames

Hey there! Let's break down how to set up the input_shape correctly for your LSTM multi-to-one architecture, since you're working with 19 sequential (128,128,3) image frames.

First, Let's Recall LSTM Input Structure

In Keras, LSTMs expect input in the format (batch_size, timesteps, data_dim)—but when defining input_shape in the layer, you omit the batch size (it's handled automatically during training).

You’re correct that each image frame has 128*128*3 = 49152 total pixels, but there are two common approaches to structure this input, depending on whether you use a standard LSTM or a more image-friendly ConvLSTM.


Approach 1: Standard LSTM (Flatten Each Frame)

If you stick with a regular LSTM, you need to flatten each 2D+channel image into a 1D vector first. This converts your input data from (batch_size, 19, 128, 128, 3) to (batch_size, 19, 49152).

Here’s how to adjust your code:

from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import LSTM, Dense

timesteps = 19
data_dim = 128 * 128 * 3  # 49152 total features per frame

model = Sequential()
# Set input_shape to (timesteps, data_dim) — no batch size needed
model.add(LSTM(units=64, input_shape=(timesteps, data_dim)))  # Adjust units based on your task
# Add your output layer (e.g., Dense(10) for 10-class classification, Dense(1) for regression)
model.add(Dense(units=1, activation='sigmoid'))  # Example for binary classification

model.compile(optimizer='adam', loss='binary_crossentropy')

Note: Flattening images loses spatial information (like edge or texture relationships), which isn’t ideal for image data. That’s why the next approach is often better.


Approach 2: ConvLSTM2D (Better for Image Sequences)

ConvLSTM2D is designed specifically for sequential image data—it preserves spatial features by applying convolutions across both space and time frames. This means you don’t need to flatten your images at all.

For a multi-to-one setup, set return_sequences=False (since we only want one final output, not outputs for each timestep):

from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import ConvLSTM2D, Flatten, Dense

model = Sequential()
# input_shape here is (timesteps, height, width, channels)
model.add(ConvLSTM2D(
    filters=32,
    kernel_size=(3, 3),
    input_shape=(19, 128, 128, 3),
    return_sequences=False  # Critical for multi-to-one output
))
# Flatten the ConvLSTM output to feed into a dense layer
model.add(Flatten())
# Add your output layer
model.add(Dense(units=1, activation='linear'))  # Example for regression

model.compile(optimizer='adam', loss='mse')

Key Takeaways

  • Standard LSTM: Use input_shape=(19, 49152) after flattening each frame.
  • ConvLSTM2D: Use input_shape=(19, 128, 128, 3) to keep spatial information intact (recommended for image sequences).

内容的提问来源于stack exchange,提问作者Yaman K Singla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:12:40