如何获取TensorFlow中MobileNet V1的最后全连接/特征层以输入LSTM
Hey there! Let's figure out which layer to pick for your LSTM input—since you've already fine-tuned MobileNet V1 for 4-class classification, you're just one step away from getting the right feature representation.
Key Layers to Consider
MobileNet V1 has two ideal candidates for feature extraction, depending on what kind of features you need for your LSTM:
Global Average Pooling Layer (
MobilenetV1/Logits/AvgPool_1a/AvgPool)- This is the go-to choice if you want a compact, fixed-dimension global feature vector. It takes the spatial output from the last convolutional block and averages it down to a 1024-dimensional vector (default for MobileNet V1, regardless of input image size).
- Perfect for sequence tasks where each frame/image needs to be represented as a single vector—you can stack these vectors into a sequence tensor (shape:
(batch_size, sequence_length, 1024)) which LSTMs handle seamlessly.
Final Convolutional Block Output (
MobilenetV1/Conv2d_13_pointwise/Relu6)- If you need richer spatial features (e.g., to model spatial patterns in your sequence), this layer is your best bet. It outputs a 7x7x1024 feature map (for 224x224 input) that retains spatial information.
- You can reshape this into a sequence by flattening the spatial dimensions (e.g.,
(batch_size, 49, 1024)where 49 = 7*7) to feed into your LSTM, giving it access to fine-grained spatial details.
Avoid These Layers
Skip the final classification layers like MobilenetV1/Logits/Conv2d_1c_1x1 (your retrained 4-class layer) or the Softmax output—these are already tailored to your 4-class task and won't provide the general-purpose features an LSTM needs.
How to Extract the Features
Here's a quick code snippet to grab your desired layer's output:
import tensorflow as tf # Load your retrained MobileNet V1 model model = tf.keras.models.load_model('your_retrained_model.h5') # First, confirm the layer names (you mentioned you already did this, but just to double-check) for layer in model.layers: print(layer.name) # Create a feature extractor for the global average pooling layer feature_extractor = tf.keras.Model( inputs=model.input, outputs=model.get_layer('MobilenetV1/Logits/AvgPool_1a/AvgPool').output ) # Or, if you want spatial features, use the final convolutional block: # feature_extractor = tf.keras.Model( # inputs=model.input, # outputs=model.get_layer('MobilenetV1/Conv2d_13_pointwise/Relu6').output # ) # Test with a sample input sample_input = tf.random.normal([1, 224, 224, 3]) # Match your input size features = feature_extractor(sample_input) print(f"Feature shape: {features.shape}") # For global pooling: (1, 1024) # For conv block: (1, 7, 7, 1024)
Quick Tip
If your LSTM expects a sequence of fixed-length vectors, the global pooling layer's output is the easiest to work with—no extra reshaping needed beyond stacking frames into a sequence. If you need spatial context, reshaping the convolutional block's output into (sequence_length, feature_dim) will do the trick.
内容的提问来源于stack exchange,提问作者Pototo

