You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python下DCGAN生成128×128图像报卷积尺寸错误的排查方案

PyTorch DCGAN 128×128图像生成卷积尺寸不匹配报错修复

错误原因

公开场景下默认的DCGAN实现均针对64×64尺寸输入设计,判别器共使用4层步长为2、卷积核尺寸4×4的下采样卷积,特征尺寸变化路径为64→32→16→8→4,最后一层4×4卷积核恰好作用在4×4特征图上,输出1×1的判别结果。
仅修改IMAGE_SIZE参数为128但不调整网络层数时,下采样到判别器最后一层前的特征图尺寸仅为2×2,无法匹配4×4的卷积核尺寸,直接触发Kernel size can't be greater than actual input size运行时错误。

修复方法

在判别器、生成器中各新增一层对应维度的卷积/转置卷积层,对齐128尺寸的上下采样路径,所有卷积参数保持DCGAN论文标准配置(卷积核4×4、下采样/上采样步长2、填充1):

  • 判别器新增一层开头下采样层,特征尺寸变化路径调整为128→64→32→16→8→4→1
  • 生成器对应新增一层末尾上采样层,特征尺寸变化路径调整为1→4→8→16→32→64→128

修正后网络代码

import torch
import torch.nn as nn

def initialize_weights(model):
    # 保留原DCGAN权重初始化逻辑
    for m in model.modules():
        if isinstance(m, (nn.Conv2d, nn.ConvTranspose2d, nn.BatchNorm2d)):
            nn.init.normal_(m.weight.data, 0.0, 0.02)

class Discriminator(nn.Module):
    def __init__(self, channels_img, features_d):
        super().__init__()
        self.disc = nn.Sequential(
            # 输入尺寸: N x 3 x 128 x 128
            nn.Conv2d(channels_img, features_d, kernel_size=4, stride=2, padding=1, bias=False),
            nn.LeakyReLU(0.2, inplace=True),
            self._block(features_d, features_d*2, 4, 2, 1),   # 64 -> 32
            self._block(features_d*2, features_d*4, 4, 2, 1), # 32 -> 16
            self._block(features_d*4, features_d*8, 4, 2, 1), # 16 -> 8
            self._block(features_d*8, features_d*16, 4, 2, 1),# 8 -> 4
            # 输出层: 4 -> 1
            nn.Conv2d(features_d*16, 1, kernel_size=4, stride=1, padding=0, bias=False),
            nn.Sigmoid()
        )

    def _block(self, in_channels, out_channels, kernel_size, stride, padding):
        return nn.Sequential(
            nn.Conv2d(in_channels, out_channels, kernel_size, stride, padding, bias=False),
            nn.BatchNorm2d(out_channels),
            nn.LeakyReLU(0.2, inplace=True)
        )

    def forward(self, x):
        return self.disc(x)

class Generator(nn.Module):
    def __init__(self, noise_dim, channels_img, features_g):
        super().__init__()
        self.gen = nn.Sequential(
            # 输入尺寸: N x 128 x 1 x 1
            nn.ConvTranspose2d(noise_dim, features_g*16, kernel_size=4, stride=1, padding=0, bias=False),
            nn.BatchNorm2d(features_g*16),
            nn.ReLU(inplace=True),
            self._block(features_g*16, features_g*8, 4, 2, 1), # 4 -> 8
            self._block(features_g*8, features_g*4, 4, 2, 1),  # 8 -> 16
            self._block(features_g*4, features_g*2, 4, 2, 1),  # 16 -> 32
            self._block(features_g*2, features_g, 4, 2, 1),    # 32 -> 64
            # 输出层: 64 -> 128
            nn.ConvTranspose2d(features_g, channels_img, kernel_size=4, stride=2, padding=1, bias=False),
            nn.Tanh()
        )

    def _block(self, in_channels, out_channels, kernel_size, stride, padding):
        return nn.Sequential(
            nn.ConvTranspose2d(in_channels, out_channels, kernel_size, stride, padding, bias=False),
            nn.BatchNorm2d(out_channels),
            nn.ReLU(inplace=True)
        )

    def forward(self, x):
        return self.gen(x)

训练注意事项

  • 原有训练超参数可直接使用,若训练初期出现模式崩塌,可将学习率调整为DCGAN论文默认的0.0002
  • 数据集加载时需严格将图像resize为128×128,像素值归一化到[-1, 1]区间,匹配生成器Tanh激活的输出范围
  • 若显存不足,可将FEATURES_DISC、FEATURES_GEN参数下调为64,仅减少卷积通道数,不影响尺寸适配逻辑

内容的提问来源于stack exchange,提问作者sid1994s

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 08:57:20