将GAN生成器输入从256x256改为32x32时遇ValueError错误求助
问题:将GAN生成器输入从256x256改为32x32后报错
错误信息
ValueError: Expected more than 1 value per channel when training, got input size torch.Size([1, 512, 1, 1])
问题原因
你使用的生成器是为256x256输入设计的,包含7次下采样操作(initial_down + 6个down Block + bottleneck)。对于32x32的输入,每次下采样步长为2,经过多次压缩后,最终瓶颈层的输出空间维度会被压缩到1x1:
- 初始输入:32x32
initial_down后:16x16down1后:8x8down2后:4x4down3后:2x2down4后:1x1
后续的down5、down6以及bottleneck仍会对1x1的特征图进行下采样,最终得到的特征图空间维度还是1x1。而BatchNorm2d在训练模式下要求每个通道至少有2个样本值来计算均值和方差,因此触发错误。
解决方案
方案1:减少下采样层数(适配32x32输入)
修改生成器,移除多余的下采样模块,让瓶颈层输入保持大于1的空间维度。针对32x32输入,只保留到down3即可,调整后的代码如下:
import torch import torch.nn as nn class Block(nn.Module): def __init__(self, in_channels, out_channels, down=True, act="relu", use_dropout=False): super(Block, self).__init__() self.conv = nn.Sequential( nn.Conv2d(in_channels, out_channels, 4, 2, 1, bias=False, padding_mode="reflect") if down else nn.ConvTranspose2d(in_channels, out_channels, 4, 2, 1, bias=False), nn.BatchNorm2d(out_channels), nn.ReLU() if act == "relu" else nn.LeakyReLU(0.2), ) self.use_dropout = use_dropout self.dropout = nn.Dropout(0.5) self.down = down def forward(self, x): x = self.conv(x) return self.dropout(x) if self.use_dropout else x class Generator(nn.Module): def __init__(self, in_channels=3, features=64): super().__init__() self.initial_down = nn.Sequential( nn.Conv2d(in_channels, features, 4, 2, 1, padding_mode="reflect"), nn.LeakyReLU(0.2), ) # 保留3次下采样,适配32x32输入 self.down1 = Block(features, features * 2, down=True, act="leaky", use_dropout=False) self.down2 = Block(features * 2, features * 4, down=True, act="leaky", use_dropout=False) self.down3 = Block(features * 4, features * 8, down=True, act="leaky", use_dropout=False) self.bottleneck = nn.Sequential( nn.Conv2d(features * 8, features * 8, 4, 2, 1), nn.ReLU() ) # 对应调整上采样层数,保持和下采样对称 self.up1 = Block(features * 8, features * 8, down=False, act="relu", use_dropout=True) self.up2 = Block(features * 8 * 2, features * 4, down=False, act="relu", use_dropout=False) self.up3 = Block(features * 4 * 2, features * 2, down=False, act="relu", use_dropout=False) self.up4 = Block(features * 2 * 2, features, down=False, act="relu", use_dropout=False) self.final_up = nn.Sequential( nn.ConvTranspose2d(features * 2, in_channels, kernel_size=4, stride=2, padding=1), nn.Tanh(), ) def forward(self, x): d1 = self.initial_down(x) d2 = self.down1(d1) d3 = self.down2(d2) d4 = self.down3(d3) bottleneck = self.bottleneck(d4) up1 = self.up1(bottleneck) up2 = self.up2(torch.cat([up1, d4], 1)) up3 = self.up3(torch.cat([up2, d3], 1)) up4 = self.up4(torch.cat([up3, d2], 1)) return self.final_up(torch.cat([up4, d1], 1)) def test(): x = torch.randn((1, 3, 32, 32)) model = Generator(in_channels=3, features=64) preds = model(x) print(preds.shape) # 输出应为torch.Size([1, 3, 32, 32]) if __name__ == "__main__": test()
方案2:测试时切换到eval模式
如果只是想快速验证模型运行,不需要训练,可以将模型设置为评估模式,此时BatchNorm会使用训练阶段统计的均值和方差,而非实时计算:
def test(): x = torch.randn((1, 3, 32, 32)) model = Generator(in_channels=3, features=64) model.eval() # 切换到eval模式 with torch.no_grad(): # 可选,减少内存占用 preds = model(x) print(preds.shape)
注意:该方法仅适合测试场景,训练时仍会触发原错误。
方案3:调整输入尺寸为适配原模型的2的幂次
原模型的下采样次数要求输入尺寸为2^8=256,如果要保留原模型结构,可以选择输入尺寸为64x64(2^6),此时瓶颈层输出为2x2,不会触发BatchNorm错误。
内容的提问来源于stack exchange,提问作者Upanshu Srivastava
相关产品推荐
相关产品推荐

