为何全局平均池化未提升Horse Or Human数据集训练速度?
Great question! The key reason your training time stays roughly the same despite using Global Average Pooling (GAP) to cut down trainable parameters is the bottleneck of your training pipeline isn't the fully connected layers—it's the frozen convolutional base. Let's break this down:
1. Convolutional layers dominate compute time
When you freeze all layers of the VGG16 base (layer.trainable = False), those convolutional layers still do all their forward pass computation every epoch—this is where almost all your training time goes.
Let's do a quick compute comparison:
- For your 160x160 input, VGG16's final convolutional layer outputs a 5x5x512 feature map.
- With
Flatten(), you turn that into a 12800-dimensional vector, then pass it to a 512-unit Dense layer: that's ~6.5 million operations. - With
GlobalAvgPool2D(), you collapse that feature map into a 512-dimensional vector, then pass it to the same Dense layer: that's ~260k operations.
- With
But compare that to just one convolutional layer in VGG16—say, the first 64-filter 3x3 conv layer: it does ~14.7 million operations per image. The deeper 512-filter conv layers do hundreds of millions of operations per batch. The fully connected/GAP portion is a tiny fraction of the total compute—so cutting that down barely moves the needle on total epoch time.
2. Frozen layers still cost time (and data loading is fixed)
Even though you're not updating the convolutional weights, you still have to run the forward pass through them to generate the feature maps for the top layers. Plus, your data generator is doing the same amount of preprocessing (augmentations, resizing) for both models—since you're using the same batch size and target size, that part of the pipeline takes identical time regardless of the top layers.
3. Why larger target sizes or other models don't help
When you scale up the target size, the convolutional layers' compute time increases even more (conv compute scales with the square of input size), making the top layers' compute an even smaller percentage of the total. The same logic applies to other base models—if you freeze their convolutional layers, the top layers' compute is negligible in comparison.
How to verify this
Try unfreezing some of the VGG16 base layers and re-running the training. You'll immediately see a jump in epoch time, because now you're doing backpropagation and weight updates for those compute-heavy convolutional layers. That's when you'll notice a difference between models with fewer trainable parameters (like the GAP version) and more (like the Flatten version)—but only when the convolutional layers are part of the training process.
内容的提问来源于stack exchange,提问作者Yoskutik

