为何全局平均池化(GAP)能够正常工作?——Network in Network相关疑问
Great question—this is such a common point of confusion when first switching from fully connected layers to GAP in CNNs. You’re totally right that theoretically, you could construct two wildly different feature maps that output the same average value. But in practice, when trained end-to-end with a classification task, GAP works incredibly well for a few key reasons:
1. It drastically reduces overfitting (compared to fully connected layers)
Fully connected layers are packed with millions of trainable parameters, which makes them prone to memorizing noise or irrelevant details in training data. GAP, on the other hand, has zero trainable parameters—all it does is compute a mean across each feature map. This forces the network to learn generalizable, semantic features instead of overfitting to pixel-level quirks.
2. It enforces semantic consistency in feature maps
When training with GAP, the network learns to map each final feature map to a specific semantic concept tied to the classification task. For example, in a dog vs. cat classifier, one final feature map might activate strongly wherever there’s a dog’s ear, another for a dog’s paw, etc.
Your edge case of "different feature maps with similar GAP outputs" doesn’t hold up here because the training process shapes those feature maps to be meaningful. If two images are both dogs, their final feature maps will activate for the same set of dog-specific concepts—even if the activation locations differ (e.g., a dog on the left vs. right side of the image). The average of those activations will still reliably signal "this is a dog" because the presence of those semantic features is what matters, not their exact spatial position.
3. It preserves spatial semantics better than fully connected layers
Fully connected layers flatten 2D feature maps into a 1D vector, which destroys spatial relationships between features. GAP keeps each feature map’s identity intact: each output node directly corresponds to the global average of one feature map, so you can easily trace back which semantic concepts drove the classification decision. This interpretability is a huge bonus, and it aligns with how CNNs are designed to learn hierarchical spatial features.
To circle back to your edge case
While you could manually create two feature maps with the same average, those maps wouldn’t be the kind the network produces during training. The model’s convolution layers are trained to generate feature maps that are highly correlated with the target classes. A random "trick" feature map that averages to the same value as a dog-specific map would never be learned, because it doesn’t help the network correctly classify images.
内容的提问来源于stack exchange,提问作者adayoegi

