You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CS231n:VGGNet总内存计算及网络结构细节技术问询

VGGNet Architecture Deep Dive

I’ve worked through the CS231n convolutional networks material extensively, so let’s break down VGGNet in a hands-on, case-based way—it’s such a foundational architecture, and its simplicity makes it perfect for understanding core convnet patterns.

Core Building Blocks

VGGNet’s design is intentionally uniform, relying on just two repeating layer types across all stages:

  • CONV Layers: Every convolutional layer uses a 3x3 filter with a stride of 1 and padding of 1. This setup guarantees the spatial dimensions of the input feature map stay identical after convolution (padding=1 offsets the edge loss from the 3x3 filter, and stride=1 means we don’t skip any pixels).
  • POOL Layers: All pooling layers are 2x2 max-pooling operations with a stride of 2 and no padding. This halves both the width and height of the feature map every time it’s applied.

Feature Map Size Evolution Walkthrough

Let’s track how dimensions change using a standard 224x224 RGB input (the default for VGGNet):

  1. Input: 224x224x3 (RGB image)
  2. First CONV Block: 2-3 stacked 3x3 CONV layers → output remains 224x224x64 (64 is the starting filter count for VGG16)
  3. First POOL Layer: 2x2 max-pool → output shrinks to 112x112x64
  4. Second CONV Block: 2-3 stacked 3x3 CONV layers (filter count doubles to 128) → stays 112x112x128
  5. Second POOL Layer: 2x2 max-pool → 56x56x128
  6. Third CONV Block: 3 stacked 3x3 CONV layers (filter count doubles to 256) → stays 56x56x256
  7. Third POOL Layer: 2x2 max-pool → 28x28x256
  8. Fourth CONV Block: 3 stacked 3x3 CONV layers (filter count doubles to 512) → stays 28x28x512
  9. Fourth POOL Layer: 2x2 max-pool → 14x14x512
  10. Fifth CONV Block: 3 stacked 3x3 CONV layers → stays 14x14x512
  11. Fifth POOL Layer: 2x2 max-pool → 7x7x512

After these conv/pool stages, the 7x7x512 feature map feeds into three fully connected (FC) layers: two FC layers with 4096 units each, followed by a final FC layer with 1000 units (for ImageNet classification).

The real genius of VGGNet is its consistency—once you grasp the "conv stack → halving pool" pattern, you can easily predict how feature maps evolve through the entire network.

内容的提问来源于stack exchange,提问作者Dang Manh Truong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:16:54