卷积神经网络(CNN)参数数量计算疑问——Andrew NG课程实例
Hey there! As someone who’s worked through Andrew Ng’s CNN course before, I totally get why this parameter count can feel confusing at first—let’s break it down clearly.
First, lock in this key rule: the number of learnable parameters in a convolutional layer only depends on your kernel (filter) size, the number of input channels, and the number of output channels (aka how many kernels you’re using). It doesn’t matter how big your input feature map’s width/height is—those dimensions don’t affect the parameter count at all!
Let’s use a typical example from Ng’s course to make this concrete:
- Suppose your input is a 3-channel image (like RGB, so input channels = 3)
- You’re using 3x3 kernels, and you have 16 of these kernels (so output channels = 16)
Here’s the step-by-step math:
- Calculate parameters per kernel: Each kernel needs to connect to every channel in the input. That’s
3 (kernel width) * 3 (kernel height) * 3 (input channels) = 27weight parameters. - Add the bias term: Every kernel has one bias parameter, so that’s +1 per kernel, making it
27 + 1 = 28parameters per kernel. - Multiply by the number of kernels:
28 * 16 = 448total learnable parameters for this convolutional layer.
Quick side note: If your example includes pooling layers (like max pooling or average pooling), those have 0 learnable parameters—they just downsample the feature map by taking max/average values, no weights or biases to train.
Let’s verify another common course example: if you have a 1-channel input (like MNIST grayscale), 5x5 kernels, and 6 output channels. The calculation would be:6 * (5*5*1 + 1) = 6 * 26 = 156 total parameters.
The core intuition here is that each kernel is reused across the entire input feature map—you don’t have separate weights for every position the kernel slides over. That’s why the input’s width/height never factors into the parameter count!
内容的提问来源于stack exchange,提问作者user9516512

