You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

胶囊网络用于FER的卷积尺寸疑问及48x48输入超参数咨询

Understanding Capsule Network Output Sizes & Adapting to 48x48 FER Inputs

Great question! Let's break this down step by step to clear up your confusion and give you practical guidance for your FER task with the fer2013 dataset.

Why the Original CapsNet Gets 6x6 from 20x20 Input

You’re right that a naive floating-point calculation of (20 - 9)/2 + 1 gives 6.5, which doesn’t make sense for a feature map. The key here is how deep learning frameworks (like TensorFlow/PyTorch) handle integer arithmetic for convolutional output sizes:

  • Frameworks use integer division (rounding down the result) when computing output dimensions, discarding any fractional part.
  • For the Primary Capsules layer in the original CapsNet, the calculation becomes:
    Output size = floor((Input size - Kernel size) / Stride) + 1
    
    Plugging in the numbers: floor((20 - 9)/2) + 1 = floor(11/2) + 1 = 5 + 1 = 6
  • No padding is used here—this integer division trick is exactly how we get the clean 6x6 feature map output from the 20x20 input.

Adapting Hyperparameters for 48x48 FER Images

You can’t directly reuse the original 28x28 MNIST hyperparameters for 48x48 FER inputs—here’s how to adjust them while preserving the CapsNet’s core logic:

1. First Convolutional Layer

The original layer uses a 9x9 kernel, stride 1, no padding to go from 28x28 → 20x20. For 48x48 inputs:

  • Keep the same kernel size and stride for consistency: (48 - 9)/1 + 1 = 40 → output is 40x40 feature maps. This works well for capturing facial details like edges and contours.
  • If you want to match the original’s relative reduction ratio (20/28 ≈ 0.71), use a 15x15 kernel: (48 -15)/1 +1 =34 (34/48≈0.71), but this isn’t strictly necessary.

2. Primary Capsules Layer

Balance feature reduction with preserving spatial information critical for facial expressions using these options:

  • Option 1: Keep original kernel/stride
    • Using 9x9 kernel, stride 2 on the 40x40 input gives: floor((40-9)/2)+1=16 → 16x16x32x8 capsules. This increases primary capsule count (from 1152 to 8192) for more capacity—offset computation by reducing primary capsule channels from 32 to 16 (keep 8 dimensions per capsule).
  • Option 2: Adjust stride for similar output size to original
    • If you want a 6x6-like spatial granularity, use stride 5: floor((40-9)/5)+1=7 →7x7x32x8 capsules. This keeps capsule count close to the original (1568 vs 1152) without major computation spikes.

3. Digit Capsules & Dynamic Routing

  • Set digit capsule count to match fer2013’s 7 expression classes (instead of the original 10 for MNIST).
  • Keep dynamic routing iterations at 3 (original value), or drop to 2 to speed up training without significant performance loss.

4. Additional Tips

  • Since fer2013 is grayscale (same as MNIST), keep input channel count at 1.
  • Add batch normalization after the first convolution to stabilize training for larger inputs.

内容的提问来源于stack exchange,提问作者nirvair

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:46:17