You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算U-Net上采样中的卷积层尺寸?以8×8到16×16为例

U-Net Upsampling (Transposed Convolution) Output Size Formula

Great question—let’s break this down clearly, since U-Net’s upsampling is a core part of its ability to recover high-resolution spatial details for segmentation tasks.

First, a quick recap: U-Net relies on transposed convolution (often loosely called "deconvolution") for upsampling. Unlike regular convolution which shrinks feature maps, transposed convolution expands them while preserving channel information.

Core Formula for Transposed Convolution Output Size

The general formula to calculate the output spatial dimensions (H_out, W_out) from input dimensions (H_in, W_in) is:

H_out = (H_in - 1) * S - 2*p + k + op
W_out = (W_in - 1) * S - 2*p + k + op

Where:

  • S: Stride of the transposed convolution (in standard U-Net, this is almost always 2 when doubling spatial size)
  • p: Padding applied to the input feature map
  • k: Size of the transposed convolution kernel
  • op: Output padding (a transposed-convolution-specific parameter to fine-tune edge dimensions)

Applying to Your Examples

Let’s use your concrete cases to make this tangible:

Example 1: 8×8 → 16×16 Feature Map

In standard U-Net implementations, the upsampling transposed convolution uses these parameters: k=4, S=2, p=1, op=0. Plugging into the formula:

H_out = (8 - 1)*2 - 2*1 + 4 + 0 = 14 - 2 + 4 = 16

This perfectly gives the 16×16 output you’re targeting.

Example 2: 8×8×32 → 16×16×32 Feature Layer

For the channel dimension: to keep the channel count at 32, the transposed convolution layer will have 32 output channels (matching the input channel count). The spatial dimension calculation is identical to Example 1—you’ll still get 16×16, and the channel dimension stays 32 because the number of transposed convolution kernels equals the input channel count.

Quick U-Net-Specific Notes

  • In most U-Net variants, the upsampled feature map is concatenated with the corresponding feature map from the downsampling path. But the size calculation above only concerns the upsampling step itself.
  • For U-Net’s symmetric encoder-decoder structure, doubling spatial size via S=2 is standard. If you needed a different output size, you could adjust S or op, but that’s rare in vanilla U-Net.

内容的提问来源于stack exchange,提问作者Sciencemoallim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 00:07:59