如何调整GAN生成器输出至28x280x3规格?
Hey there! Let's get your generator to output that 28x280x3 image when face_sequence=True. I spot a few issues in your current code that are causing errors, plus a couple of straightforward ways to hit your target dimensions.
First off, those Reshape and Lambda layers you added are breaking the spatial structure of your feature maps and causing dimension mismatches — we can ditch those entirely. Here are two solid solutions:
Solution 1: Use the Resizing Layer (Simplest Fix)
This is the quickest way to adjust your existing pipeline without reworking all your upsampling steps. The Resizing layer lets you directly stretch or shrink your feature map to the exact dimensions you need.
from tensorflow.keras.layers import Resizing def build_generator_face(latent_dim, channels, face_sequence): model = Sequential() # Fixed input_shape: your noise is a 1D vector, no need for the extra "None" dimension model.add(Dense(128 * 7 * 7, activation="relu", input_shape=(latent_dim,))) model.add(Reshape((7, 7, 128))) model.add(UpSampling2D()) # Default (2,2) → 14x14x128 model.add(Conv2D(128, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) model.add(UpSampling2D()) # → 28x28x128 model.add(Conv2D(64, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) if face_sequence == False: model.add(Conv2D(64, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) model.add(Conv2D(64, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) else: # Three (1,2) upsamplings get us to 28x224x64 model.add(UpSampling2D(size=(1, 2))) # 28x56x64 model.add(Conv2D(64, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) model.add(UpSampling2D(size=(1, 2))) # 28x112x64 model.add(Conv2D(64, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) model.add(UpSampling2D(size=(1, 2))) # 28x224x64 model.add(Conv2D(64, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) # Generate 3-channel features first model.add(Conv2D(channels, kernel_size=4, padding="same")) model.add(Activation("relu")) # Keep activation linear before resizing for better interpolation # The magic: resize width from 224 to 280, keep height at 28 model.add(Resizing(height=28, width=280, interpolation='bilinear')) # Apply final tanh activation for proper output range model.add(Activation("tanh")) model.summary() noise = Input(shape=(latent_dim,)) img = model(noise) mdl = Model(noise, img) return mdl
Key Fixes & Explanations:
- Removed all broken
Reshape/Lambdalayers — these were mangling your feature map dimensions and causing errors. - Added
Resizingto stretch the 28x224x3 output to 28x280x3. Bilinear interpolation ensures smooth resizing without jagged edges. - Fixed the
input_shapefor your initialDenselayer — your noise vector is 1D, so(latent_dim,)is correct (no extraNone). - Reordered activations: Using
relubefore resizing avoids clamping values to the tanh range during interpolation, which helps preserve detail.
Solution 2: Adjust Upsampling with Conv2DTranspose (GAN-Friendly)
If you prefer to avoid resizing and want to generate the exact dimension via convolution, you can replace your final upsampling step with Conv2DTranspose. Note this works best for integer-scale upsizes, but we can tweak it for your needs:
from tensorflow.keras.layers import Conv2DTranspose def build_generator_face(latent_dim, channels, face_sequence): model = Sequential() model.add(Dense(128 * 7 * 7, activation="relu", input_shape=(latent_dim,))) model.add(Reshape((7, 7, 128))) model.add(UpSampling2D()) model.add(Conv2D(128, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) model.add(UpSampling2D()) model.add(Conv2D(64, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) if face_sequence == False: model.add(Conv2D(64, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) model.add(Conv2D(64, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) else: model.add(UpSampling2D(size=(1, 2))) model.add(Conv2D(64, kernel_size=4, padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) # Replace the next two upsamplings with a single Conv2DTranspose to jump to 280 width # 28x56 → 28x280 via stride=(1,5) model.add(Conv2DTranspose(64, kernel_size=(4, 10), strides=(1, 5), padding="same")) model.add(BatchNormalization(momentum=0.8)) model.add(Activation("relu")) model.add(Conv2D(channels, kernel_size=4, padding="same")) model.add(Activation("tanh")) model.summary() noise = Input(shape=(latent_dim,)) img = model(noise) mdl = Model(noise, img) return mdl
Note:
Conv2DTransposeuses learnable filters to upsample, which can help preserve better spatial details for GANs. However, since 224 to 280 isn't an integer power of 2, we use a stride of (1,5) to jump from 56 to 280 directly.
Verify Your Output
After making these changes, run model.summary() — when face_sequence=True, the final output shape should be (None, 28, 280, 3) (the None is your batch dimension, which is expected).
内容的提问来源于stack exchange,提问作者Jose Ramon

