You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ResNet维度提升方法探究:架构中虚线是否为卷积升维层?

How ResNet Handles Dimension Expansion (And That Dashed Connection You’re Curious About)

Great question—ResNet’s ability to adjust feature dimensions while preserving residual flow is one of its key innovations. Let’s break this down clearly, especially focusing on that dashed connection in the architecture diagrams.

Why ResNet Needs Dimension Expansion

As the network deepens, we need to:

  • Increase the number of feature channels to capture more complex, abstract features (e.g., going from 64 to 128 channels)
  • Sometimes reduce the spatial size of feature maps (downsampling) to cut computation costs and focus on global patterns

The residual shortcut path has to match the dimension of the main path’s output to enable element-wise addition. Dashed lines indicate when this shortcut needs to modify the input dimensions to make that addition possible.

Two Main Ways ResNet Achieves Dimension Expansion

1. 1x1 Convolution with Stride (Most Common in Deeper ResNets)

This is the standard approach you’ll see in dashed connections for models like ResNet-50/101. Here’s how it works:

  • Channel increase: A 1x1 convolution layer uses enough kernels to match the target channel count. For example, if the input is 64 channels, we use 128 1x1 kernels to bump it to 128 channels.
  • Spatial downsampling: The convolution uses a stride of 2, which halves the width and height of the feature map (e.g., 56x56 → 28x28).
  • No ReLU here: Unlike the main path, this shortcut convolution doesn’t include a ReLU activation—we want to preserve the original input’s information flow without distorting it before adding to the main path.

This operation perfectly aligns the shortcut’s output with the main path (which also uses stride=2 in its 3x3 convolution for downsampling) so they can be added together.

2. Zero Padding (Used in Shallow ResNets Like 18/34)

For cases where we only need to increase channels without downsampling, ResNet uses channel-wise zero padding:

  • We append zero-filled channels to the input feature map. For example, a 64-channel input gets 64 zero channels added, making it 128 channels total.
  • This is cheaper computationally than a convolution, but it doesn’t learn any new features—just fills the dimension gap. It’s mostly used in the earlier, smaller ResNet variants.

Is the Dashed Line Only a Conv Layer for Dimension Increase?

Short answer: No—it’s a dimension-matching module, and its implementation depends on what’s needed:

  • If we need both downsampling and channel increase: The dashed connection is a 1x1 convolution with stride=2 (plus Batch Normalization, no ReLU).
  • If we only need channel increase (no spatial change): The dashed connection is just zero padding, no convolution involved.

In most standard architecture diagrams (especially for deeper ResNets), the dashed line refers to the 1x1 convolution approach because that’s the more common scenario as the network scales up.

内容的提问来源于stack exchange,提问作者Sudip Das

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:48:05