You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整GAN生成器输出至28x280x3规格?

Hey there! Let's get your generator to output that 28x280x3 image when face_sequence=True. I spot a few issues in your current code that are causing errors, plus a couple of straightforward ways to hit your target dimensions.

First off, those Reshape and Lambda layers you added are breaking the spatial structure of your feature maps and causing dimension mismatches — we can ditch those entirely. Here are two solid solutions:


Solution 1: Use the Resizing Layer (Simplest Fix)

This is the quickest way to adjust your existing pipeline without reworking all your upsampling steps. The Resizing layer lets you directly stretch or shrink your feature map to the exact dimensions you need.

from tensorflow.keras.layers import Resizing

def build_generator_face(latent_dim, channels, face_sequence):
    model = Sequential()
    # Fixed input_shape: your noise is a 1D vector, no need for the extra "None" dimension
    model.add(Dense(128 * 7 * 7, activation="relu", input_shape=(latent_dim,)))
    model.add(Reshape((7, 7, 128)))
    model.add(UpSampling2D())  # Default (2,2) → 14x14x128
    model.add(Conv2D(128, kernel_size=4, padding="same"))
    model.add(BatchNormalization(momentum=0.8))
    model.add(Activation("relu"))
    model.add(UpSampling2D())  # → 28x28x128
    model.add(Conv2D(64, kernel_size=4, padding="same"))
    model.add(BatchNormalization(momentum=0.8))
    model.add(Activation("relu"))
    
    if face_sequence == False:
        model.add(Conv2D(64, kernel_size=4, padding="same"))
        model.add(BatchNormalization(momentum=0.8))
        model.add(Activation("relu"))
        model.add(Conv2D(64, kernel_size=4, padding="same"))
        model.add(BatchNormalization(momentum=0.8))
        model.add(Activation("relu"))
    else:
        # Three (1,2) upsamplings get us to 28x224x64
        model.add(UpSampling2D(size=(1, 2)))  # 28x56x64
        model.add(Conv2D(64, kernel_size=4, padding="same"))
        model.add(BatchNormalization(momentum=0.8))
        model.add(Activation("relu"))
        model.add(UpSampling2D(size=(1, 2)))  # 28x112x64
        model.add(Conv2D(64, kernel_size=4, padding="same"))
        model.add(BatchNormalization(momentum=0.8))
        model.add(Activation("relu"))
        model.add(UpSampling2D(size=(1, 2)))  # 28x224x64
        model.add(Conv2D(64, kernel_size=4, padding="same"))
        model.add(BatchNormalization(momentum=0.8))
        model.add(Activation("relu"))
    
    # Generate 3-channel features first
    model.add(Conv2D(channels, kernel_size=4, padding="same"))
    model.add(Activation("relu"))  # Keep activation linear before resizing for better interpolation
    
    # The magic: resize width from 224 to 280, keep height at 28
    model.add(Resizing(height=28, width=280, interpolation='bilinear'))
    
    # Apply final tanh activation for proper output range
    model.add(Activation("tanh"))
    
    model.summary()
    noise = Input(shape=(latent_dim,))
    img = model(noise)
    mdl = Model(noise, img)
    return mdl

Key Fixes & Explanations:

  • Removed all broken Reshape/Lambda layers — these were mangling your feature map dimensions and causing errors.
  • Added Resizing to stretch the 28x224x3 output to 28x280x3. Bilinear interpolation ensures smooth resizing without jagged edges.
  • Fixed the input_shape for your initial Dense layer — your noise vector is 1D, so (latent_dim,) is correct (no extra None).
  • Reordered activations: Using relu before resizing avoids clamping values to the tanh range during interpolation, which helps preserve detail.

Solution 2: Adjust Upsampling with Conv2DTranspose (GAN-Friendly)

If you prefer to avoid resizing and want to generate the exact dimension via convolution, you can replace your final upsampling step with Conv2DTranspose. Note this works best for integer-scale upsizes, but we can tweak it for your needs:

from tensorflow.keras.layers import Conv2DTranspose

def build_generator_face(latent_dim, channels, face_sequence):
    model = Sequential()
    model.add(Dense(128 * 7 * 7, activation="relu", input_shape=(latent_dim,)))
    model.add(Reshape((7, 7, 128)))
    model.add(UpSampling2D())
    model.add(Conv2D(128, kernel_size=4, padding="same"))
    model.add(BatchNormalization(momentum=0.8))
    model.add(Activation("relu"))
    model.add(UpSampling2D())
    model.add(Conv2D(64, kernel_size=4, padding="same"))
    model.add(BatchNormalization(momentum=0.8))
    model.add(Activation("relu"))
    
    if face_sequence == False:
        model.add(Conv2D(64, kernel_size=4, padding="same"))
        model.add(BatchNormalization(momentum=0.8))
        model.add(Activation("relu"))
        model.add(Conv2D(64, kernel_size=4, padding="same"))
        model.add(BatchNormalization(momentum=0.8))
        model.add(Activation("relu"))
    else:
        model.add(UpSampling2D(size=(1, 2)))
        model.add(Conv2D(64, kernel_size=4, padding="same"))
        model.add(BatchNormalization(momentum=0.8))
        model.add(Activation("relu"))
        # Replace the next two upsamplings with a single Conv2DTranspose to jump to 280 width
        # 28x56 → 28x280 via stride=(1,5)
        model.add(Conv2DTranspose(64, kernel_size=(4, 10), strides=(1, 5), padding="same"))
        model.add(BatchNormalization(momentum=0.8))
        model.add(Activation("relu"))
    
    model.add(Conv2D(channels, kernel_size=4, padding="same"))
    model.add(Activation("tanh"))
    
    model.summary()
    noise = Input(shape=(latent_dim,))
    img = model(noise)
    mdl = Model(noise, img)
    return mdl

Note:

  • Conv2DTranspose uses learnable filters to upsample, which can help preserve better spatial details for GANs. However, since 224 to 280 isn't an integer power of 2, we use a stride of (1,5) to jump from 56 to 280 directly.

Verify Your Output

After making these changes, run model.summary() — when face_sequence=True, the final output shape should be (None, 28, 280, 3) (the None is your batch dimension, which is expected).

内容的提问来源于stack exchange,提问作者Jose Ramon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 12:57:41