TensorFlow Slim中AlexNetV2用conv2d替代全连接层的维度疑问求解
Hey there, let's clear up this confusion about replacing fully connected layers with conv2d in AlexNet v2!
First, let's recap the core of your question: your pool5 output is 6x6x256, and after a 5x5 conv2d (no padding, stride 1), you get 2x2x4096—but the original fully connected layer (fc6) has only 4096 neurons. Why the discrepancy?
1. The "Strict Equivalent" Conv2D for FC Layers
The original fc6 in AlexNet takes the flattened 6x6x256 tensor (total 9216 elements) and outputs a 4096-dimensional vector. To directly replicate this with a conv2d, you need to use a convolution kernel that matches the spatial size of pool5:
- Kernel size:
6x6(same as pool5's width/height) - Padding:
VALID - Filters: 4096
This would produce a 1x1x4096 tensor—squeeze the spatial dimensions, and you get exactly the same 4096-neuron output as the original fc6. The parameter count is identical too:
- FC layer params:
9216 * 4096 + 4096 - Conv2D params:
6*6*256 *4096 +4096(since6*6*256 =9216)
Here's how that would look in your code:
# Strictly equivalent conv2d replacement for fc6 net = slim.conv2d(net, 4096, [6,6], stride=1, padding='VALID', scope='fc6_conv') # Remove spatial dimensions to match original fc6 shape net = tf.squeeze(net, axis=[1,2], name='fc6_squeeze')
2. Why Your 5x5 Conv2D Produces 2x2x4096
Your current setup uses a 5x5 kernel, which is a deliberate choice that deviates from the strict fc replacement. This design has two key purposes:
- Preserve spatial information: Unlike fully connected layers (which flatten all spatial data into a 1D vector), the
2x2x4096output keeps spatial context. This is useful for tasks like object detection or semantic segmentation, where location matters. - Enable variable input sizes: Fully connected layers require fixed input dimensions (since flattening depends on spatial size). Conv2D layers let your network process larger/smaller inputs—if you feed a bigger image, pool5's spatial size grows, and the 5x5 conv will produce a larger spatial output instead of breaking.
3. Converting 2x2x4096 to 4096 Neurons
If you want to map this 2x2x4096 tensor back to the original fc6's 4096-dimensional output, you have two simple options:
- Global Average Pooling (GAP): Average the values across the
2x2spatial dimensions to get a1x1x4096tensor:net = slim.conv2d(net, 4096, [5,5], stride=1, padding='VALID', scope='fc6_conv') # Apply 2x2 average pooling to collapse spatial dimensions net = slim.avg_pool2d(net, [2,2], stride=1, padding='VALID', scope='fc6_pool') net = tf.squeeze(net, axis=[1,2], name='fc6_squeeze') - Flatten + Optional Reduction: Flatten the
2x2x4096into a 16384-dimensional vector, then use a 1x1 conv or small fc layer to reduce it to 4096 (though this undermines the conv layer's variable-size advantage).
Key Takeaway
Your 5x5 conv2d setup isn't a mistake—it's a flexible alternative to the rigid fully connected layer. The 2x2x4096 output is intentional if you want to keep spatial data; if you need exact parity with the original AlexNet's fc6, switch to a 6x6 kernel instead.
内容的提问来源于stack exchange,提问作者jing wang

