ResNet维度提升方法探究:架构中虚线是否为卷积升维层?
Great question—ResNet’s ability to adjust feature dimensions while preserving residual flow is one of its key innovations. Let’s break this down clearly, especially focusing on that dashed connection in the architecture diagrams.
Why ResNet Needs Dimension Expansion
As the network deepens, we need to:
- Increase the number of feature channels to capture more complex, abstract features (e.g., going from 64 to 128 channels)
- Sometimes reduce the spatial size of feature maps (downsampling) to cut computation costs and focus on global patterns
The residual shortcut path has to match the dimension of the main path’s output to enable element-wise addition. Dashed lines indicate when this shortcut needs to modify the input dimensions to make that addition possible.
Two Main Ways ResNet Achieves Dimension Expansion
1. 1x1 Convolution with Stride (Most Common in Deeper ResNets)
This is the standard approach you’ll see in dashed connections for models like ResNet-50/101. Here’s how it works:
- Channel increase: A 1x1 convolution layer uses enough kernels to match the target channel count. For example, if the input is 64 channels, we use 128 1x1 kernels to bump it to 128 channels.
- Spatial downsampling: The convolution uses a stride of 2, which halves the width and height of the feature map (e.g., 56x56 → 28x28).
- No ReLU here: Unlike the main path, this shortcut convolution doesn’t include a ReLU activation—we want to preserve the original input’s information flow without distorting it before adding to the main path.
This operation perfectly aligns the shortcut’s output with the main path (which also uses stride=2 in its 3x3 convolution for downsampling) so they can be added together.
2. Zero Padding (Used in Shallow ResNets Like 18/34)
For cases where we only need to increase channels without downsampling, ResNet uses channel-wise zero padding:
- We append zero-filled channels to the input feature map. For example, a 64-channel input gets 64 zero channels added, making it 128 channels total.
- This is cheaper computationally than a convolution, but it doesn’t learn any new features—just fills the dimension gap. It’s mostly used in the earlier, smaller ResNet variants.
Is the Dashed Line Only a Conv Layer for Dimension Increase?
Short answer: No—it’s a dimension-matching module, and its implementation depends on what’s needed:
- If we need both downsampling and channel increase: The dashed connection is a 1x1 convolution with stride=2 (plus Batch Normalization, no ReLU).
- If we only need channel increase (no spatial change): The dashed connection is just zero padding, no convolution involved.
In most standard architecture diagrams (especially for deeper ResNets), the dashed line refers to the 1x1 convolution approach because that’s the more common scenario as the network scales up.
内容的提问来源于stack exchange,提问作者Sudip Das

