运行Deep Convolutional Autoencoder时出现RuntimeError问题求助
卷积自编码器运行报错解决
问题场景
尝试构建解码器架构与DCGAN生成器一致的深度卷积自编码器,运行时触发维度不匹配错误。核心模型代码如下:
class Autoencoder(nn.Module): def __init__(self): super(Autoencoder, self).__init__() self.encoder = nn.Sequential( nn.Conv2d(1, 16, 3, 2, 1, bias=False), nn.LeakyReLU(0.2, inplace=True), nn.Conv2d(16, 16 * 2, 3, 2, 1, bias=False), nn.BatchNorm2d(16 * 2), nn.LeakyReLU(0.2, inplace=True), nn.Conv2d(16 * 2, 16 * 4, 3, 2, 1, bias=False), nn.BatchNorm2d(16 * 4), nn.LeakyReLU(0.2, inplace=True), nn.Conv2d(16 * 4, 16 * 8, 3, 2, 1, bias=False), nn.BatchNorm2d(16 * 8), nn.LeakyReLU(0.2, inplace=True), nn.Conv2d(16 * 8, 16 * 16, 3), nn.Sigmoid() ) self.decoder = nn.Sequential( nn.ConvTranspose2d( 16 * 16, 16 * 8, 3), nn.BatchNorm2d(64 * 8), nn.ReLU(True), nn.ConvTranspose2d(16 * 8, 16 * 4, 3, 2, 1, output_padding=1), nn.BatchNorm2d(16 * 4), nn.ReLU(True), nn.ConvTranspose2d(16 * 4, 16 * 2, 3, 2, 1, output_padding=1), nn.BatchNorm2d(16 * 2), nn.ReLU(True), nn.ConvTranspose2d(16 * 2, 16, 3, 2, 1, output_padding=1), nn.BatchNorm2d(16), nn.ReLU(True), nn.ConvTranspose2d( 16, 1, 3, 2, 1, output_padding=1), nn.Tanh() ) def forward(self, x): x = self.encoder(x) x = self.decoder(x) return x
错误信息
RuntimeError: Calculated padded input size per channel: (2 x 2). Kernel size: (3 x 3). Kernel size can't be greater than actual input size
问题分析
- 编码器最后一层卷积尺寸不匹配:以MNIST 28x28输入为例,经过前4次步长为2的卷积后,特征图尺寸变为2x2。此时最后一层
nn.Conv2d(16 * 8, 16 * 16, 3)使用3x3卷积核且无padding,输入2x2的特征图无法容纳3x3的卷积核,直接触发报错。 - 解码器BatchNorm通道数错误:
nn.BatchNorm2d(64 * 8)的通道数应为168=128,而非648=512,会导致后续维度不匹配。
修复方案
1. 修正编码器最后一层卷积
给最后一层卷积添加padding=1,保证输入2x2的特征图经过padding后可适配3x3卷积核,输出尺寸保持2x2:
nn.Conv2d(16 * 8, 16 * 16, 3, padding=1),
2. 修正解码器BatchNorm通道数
将错误的通道数改为16*8:
nn.BatchNorm2d(16 * 8),
修复后的完整模型代码
class Autoencoder(nn.Module): def __init__(self): super(Autoencoder, self).__init__() self.encoder = nn.Sequential( nn.Conv2d(1, 16, 3, 2, 1, bias=False), nn.LeakyReLU(0.2, inplace=True), nn.Conv2d(16, 16 * 2, 3, 2, 1, bias=False), nn.BatchNorm2d(16 * 2), nn.LeakyReLU(0.2, inplace=True), nn.Conv2d(16 * 2, 16 * 4, 3, 2, 1, bias=False), nn.BatchNorm2d(16 * 4), nn.LeakyReLU(0.2, inplace=True), nn.Conv2d(16 * 4, 16 * 8, 3, 2, 1, bias=False), nn.BatchNorm2d(16 * 8), nn.LeakyReLU(0.2, inplace=True), nn.Conv2d(16 * 8, 16 * 16, 3, padding=1), nn.Sigmoid() ) self.decoder = nn.Sequential( nn.ConvTranspose2d(16 * 16, 16 * 8, 3), nn.BatchNorm2d(16 * 8), nn.ReLU(True), nn.ConvTranspose2d(16 * 8, 16 * 4, 3, 2, 1, output_padding=1), nn.BatchNorm2d(16 * 4), nn.ReLU(True), nn.ConvTranspose2d(16 * 4, 16 * 2, 3, 2, 1, output_padding=1), nn.BatchNorm2d(16 * 2), nn.ReLU(True), nn.ConvTranspose2d(16 * 2, 16, 3, 2, 1, output_padding=1), nn.BatchNorm2d(16), nn.ReLU(True), nn.ConvTranspose2d(16, 1, 3, 2, 1, output_padding=1), nn.Tanh() ) def forward(self, x): x = self.encoder(x) x = self.decoder(x) return x
内容的提问来源于stack exchange,提问作者al. ekrami
相关产品推荐
相关产品推荐

