迁移学习下Keras LSTM层输入维度不兼容问题求助
Alright, let's break down exactly what's going wrong here and walk through the fixes step by step.
First, let's recap the core requirement for Keras LSTM layers: they expect 3-dimensional input with the shape (number_of_samples, timesteps, features_per_timestep). That's why you first got the error saying "expected ndim=3, found ndim=2"—your initial input (after flattening 5×5×122880 into a single vector) was 2D: (1032, 5*5*122880).
Why reshaping to (1032,25,122880) still didn't work
You were on the right track by reshaping to 3D, but there are two likely issues here:
Mismatched input shape in your model definition
If you didn't update your model's input layer to match the new 3D shape, Keras will still expect the old dimension. For example, if your input layer was defined asInput(shape=(5*5*122880,))(2D), it won't accept the 3D(25,122880)input.Fix this by explicitly defining the input shape to match your reshaped data:
from keras.layers import Input, LSTM, Dense from keras.models import Model # Input shape should be (timesteps, features) — no need to include sample count input_layer = Input(shape=(25, 122880)) # Add LSTM layer (adjust units based on your needs) lstm_layer = LSTM(units=64)(input_layer) # Final dense layer for binary classification output_layer = Dense(1, activation='sigmoid')(lstm_layer) model = Model(inputs=input_layer, outputs=output_layer) model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])That 122880 feature dimension is extremely unusual (and problematic)
Let's be real—122k features per timestep is way larger than any standard transfer learning bottleneck feature. Typical pre-trained models (like ResNet, VGG) output bottleneck features with channel counts in the hundreds or low thousands (e.g., ResNet50 outputs(samples,7,7,2048)).This giant dimension is almost certainly a mistake in how you extracted your bottleneck features:
- Did you accidentally concatenate features multiple times?
- Did you mix up batch and channel dimensions during extraction?
- Double-check your feature extraction code. For reference, here's how you'd correctly get bottleneck features from a pre-trained model:
from keras.applications.resnet50 import ResNet50, preprocess_input from keras.preprocessing.image import ImageDataGenerator base_model = ResNet50(weights='imagenet', include_top=False, input_shape=(224,224,3)) # Assume you have a generator for your training images train_datagen = ImageDataGenerator(preprocessing_function=preprocess_input) train_generator = train_datagen.flow_from_directory( 'train_dir', target_size=(224,224), batch_size=32, class_mode='binary', shuffle=False ) bottleneck_features = base_model.predict(train_generator) print(bottleneck_features.shape) # Should look like (1032,7,7,2048) — not 122880 channels
If for some reason you do need to work with 122880 features, you must add a dimension-reduction step first—otherwise your LSTM will have an impossible number of parameters (think billions) and will never train. Use a
TimeDistributedDense layer to compress each timestep's features:from keras.layers import TimeDistributed input_layer = Input(shape=(25, 122880)) # Reduce features per timestep to a manageable size (e.g., 256) reduced_features = TimeDistributed(Dense(256, activation='relu'))(input_layer) lstm_layer = LSTM(64)(reduced_features) output_layer = Dense(1, activation='sigmoid')(lstm_layer)
Final Checks
Before training, always verify your data shape matches the model's input shape:
print(train_bottleneck_features.shape) # Should output (1032,25,122880) (or your reduced feature size)
内容的提问来源于stack exchange,提问作者Rick Lentz

