在Keras中CNN后接LSTM/卷积LSTM的时序图像回归问题求助
Hey there! Let's figure out why your CNN-LSTM setup is throwing errors and fix it properly. The core issue is a mismatch between what your CNN outputs and what LSTM expects as input, plus some wrong layer ordering. Let's break this down:
Key Problems in Your Current Code
- LSTM Input Shape Mismatch: LSTMs require input in the format
(batch_size, timesteps, features), but your CNN outputs a 4D tensor of shape(batch_size, channels, height, width)(channels-first as per your input setup). Feeding this directly to LSTM will cause dimension errors immediately. - Incorrect Layer Order: You added
Flatten()after LSTM, which is backwards. You need to reshape CNN's spatial features into a sequence format before passing to LSTM, then process the LSTM output with dense layers. - Missing Temporal Context: Your input
Xis(7000, 3, 128, 128)—this suggests each sample is a single RGB image, not a sequence of images. If you're working with temporal image data, each sample should have a time dimension (e.g.,(7000, T, 3, 128, 128)whereTis the number of frames per sequence).
Fixed Code Solutions
Solution 1: If You Have Sequential Temporal Images (Each Sample Has T Frames)
First, adjust your input data to shape (7000, T, 3, 128, 128) where T is your actual number of time steps per sample. Then use TimeDistributed to apply the CNN to each frame in the sequence:
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import ( Convolution2D, Activation, MaxPooling2D, Dropout, TimeDistributed, LSTM, Flatten, Dense ) size = 128 T = 10 # Replace with your real number of time steps per sample model = Sequential() # Wrap CNN layers in TimeDistributed to process each frame independently model.add(TimeDistributed(Convolution2D(32, (3, 3)), input_shape=(T, 3, size, size))) model.add(TimeDistributed(Activation('relu'))) model.add(TimeDistributed(Convolution2D(32, (3, 3)))) model.add(TimeDistributed(Activation('relu'))) model.add(TimeDistributed(MaxPooling2D(pool_size=(2, 2)))) model.add(TimeDistributed(Dropout(0.25))) model.add(TimeDistributed(Convolution2D(64, (3, 3)))) model.add(TimeDistributed(Activation('relu'))) model.add(TimeDistributed(Convolution2D(64, (3, 3)))) model.add(TimeDistributed(Activation('relu'))) model.add(TimeDistributed(MaxPooling2D(pool_size=(2, 2)))) model.add(TimeDistributed(Dropout(0.25))) # Flatten each frame's features into a 1D vector model.add(TimeDistributed(Flatten())) # Now input to LSTM is (batch_size, T, flattened_features) which matches its requirement model.add(LSTM(50)) # Final dense layers for regression (Y ranges from 0-1) model.add(Dense(32)) model.add(Activation('relu')) model.add(Dropout(0.5)) model.add(Dense(1)) model.add(Activation('sigmoid')) model.compile(loss='mse', optimizer='adam') return model
Solution 2: If You're Treating Spatial Features as "Temporal" (Single Frame per Sample)
If you only have single frames but want to use LSTM by treating spatial positions as time steps, reshape the CNN's output into a valid sequence format:
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import ( Convolution2D, Activation, MaxPooling2D, Dropout, LSTM, Flatten, Dense, Reshape ) size = 128 model = Sequential() # CNN part (channels-first input as per your original code) model.add(Convolution2D(32, (3, 3), input_shape=(3, size, size))) model.add(Activation('relu')) model.add(Convolution2D(32, (3, 3))) model.add(Activation('relu')) model.add(MaxPooling2D(pool_size=(2, 2))) model.add(Dropout(0.25)) model.add(Convolution2D(64, (3, 3))) model.add(Activation('relu')) model.add(Convolution2D(64, (3, 3))) model.add(Activation('relu')) model.add(MaxPooling2D(pool_size=(2, 2))) model.add(Dropout(0.25)) # Calculate CNN output shape: after pooling, we get (64, 29, 29) # Reshape to (29*29, 64) → treat each spatial position as a time step, channels as features model.add(Reshape((29*29, 64))) # Now input to LSTM is (batch_size, 841, 64) which is valid model.add(LSTM(50)) # Final dense layers model.add(Dense(32)) model.add(Activation('relu')) model.add(Dropout(0.5)) model.add(Dense(1)) model.add(Activation('sigmoid')) model.compile(loss='mse', optimizer='adam') return model
Explanation
- TimeDistributed Layer: In Solution 1, this ensures the CNN processes each frame in the temporal sequence independently, preserving the time dimension for the LSTM to learn temporal patterns.
- Reshape Layer: In Solution 2, we convert the 3D spatial feature map into a 2D sequence where each "time step" is a spatial pixel's feature vector, letting the LSTM process spatial relationships as sequential data.
- Layer Order: We always reshape/flatten the CNN output before passing to LSTM, then use the LSTM's condensed sequence output to feed into dense layers for regression.
内容的提问来源于stack exchange,提问作者Yooooung
相关产品推荐
相关产品推荐

