Python下DCGAN生成128×128图像报卷积尺寸错误的排查方案
PyTorch DCGAN 128×128图像生成卷积尺寸不匹配报错修复
错误原因
公开场景下默认的DCGAN实现均针对64×64尺寸输入设计,判别器共使用4层步长为2、卷积核尺寸4×4的下采样卷积,特征尺寸变化路径为64→32→16→8→4,最后一层4×4卷积核恰好作用在4×4特征图上,输出1×1的判别结果。
仅修改IMAGE_SIZE参数为128但不调整网络层数时,下采样到判别器最后一层前的特征图尺寸仅为2×2,无法匹配4×4的卷积核尺寸,直接触发Kernel size can't be greater than actual input size运行时错误。
修复方法
在判别器、生成器中各新增一层对应维度的卷积/转置卷积层,对齐128尺寸的上下采样路径,所有卷积参数保持DCGAN论文标准配置(卷积核4×4、下采样/上采样步长2、填充1):
- 判别器新增一层开头下采样层,特征尺寸变化路径调整为
128→64→32→16→8→4→1 - 生成器对应新增一层末尾上采样层,特征尺寸变化路径调整为
1→4→8→16→32→64→128
修正后网络代码
import torch import torch.nn as nn def initialize_weights(model): # 保留原DCGAN权重初始化逻辑 for m in model.modules(): if isinstance(m, (nn.Conv2d, nn.ConvTranspose2d, nn.BatchNorm2d)): nn.init.normal_(m.weight.data, 0.0, 0.02) class Discriminator(nn.Module): def __init__(self, channels_img, features_d): super().__init__() self.disc = nn.Sequential( # 输入尺寸: N x 3 x 128 x 128 nn.Conv2d(channels_img, features_d, kernel_size=4, stride=2, padding=1, bias=False), nn.LeakyReLU(0.2, inplace=True), self._block(features_d, features_d*2, 4, 2, 1), # 64 -> 32 self._block(features_d*2, features_d*4, 4, 2, 1), # 32 -> 16 self._block(features_d*4, features_d*8, 4, 2, 1), # 16 -> 8 self._block(features_d*8, features_d*16, 4, 2, 1),# 8 -> 4 # 输出层: 4 -> 1 nn.Conv2d(features_d*16, 1, kernel_size=4, stride=1, padding=0, bias=False), nn.Sigmoid() ) def _block(self, in_channels, out_channels, kernel_size, stride, padding): return nn.Sequential( nn.Conv2d(in_channels, out_channels, kernel_size, stride, padding, bias=False), nn.BatchNorm2d(out_channels), nn.LeakyReLU(0.2, inplace=True) ) def forward(self, x): return self.disc(x) class Generator(nn.Module): def __init__(self, noise_dim, channels_img, features_g): super().__init__() self.gen = nn.Sequential( # 输入尺寸: N x 128 x 1 x 1 nn.ConvTranspose2d(noise_dim, features_g*16, kernel_size=4, stride=1, padding=0, bias=False), nn.BatchNorm2d(features_g*16), nn.ReLU(inplace=True), self._block(features_g*16, features_g*8, 4, 2, 1), # 4 -> 8 self._block(features_g*8, features_g*4, 4, 2, 1), # 8 -> 16 self._block(features_g*4, features_g*2, 4, 2, 1), # 16 -> 32 self._block(features_g*2, features_g, 4, 2, 1), # 32 -> 64 # 输出层: 64 -> 128 nn.ConvTranspose2d(features_g, channels_img, kernel_size=4, stride=2, padding=1, bias=False), nn.Tanh() ) def _block(self, in_channels, out_channels, kernel_size, stride, padding): return nn.Sequential( nn.ConvTranspose2d(in_channels, out_channels, kernel_size, stride, padding, bias=False), nn.BatchNorm2d(out_channels), nn.ReLU(inplace=True) ) def forward(self, x): return self.gen(x)
训练注意事项
- 原有训练超参数可直接使用,若训练初期出现模式崩塌,可将学习率调整为DCGAN论文默认的0.0002
- 数据集加载时需严格将图像resize为128×128,像素值归一化到[-1, 1]区间,匹配生成器Tanh激活的输出范围
- 若显存不足,可将
FEATURES_DISC、FEATURES_GEN参数下调为64,仅减少卷积通道数,不影响尺寸适配逻辑
内容的提问来源于stack exchange,提问作者sid1994s
相关产品推荐
相关产品推荐

