TensorFlow训练DCGAN遇变量不存在错误,设置reuse=True仍未解决
Hey there! Let's break down that frustrating ValueError you're hitting with your DCGAN training. The error tells us that TensorFlow can't find an EMA variable for your discriminator's batch normalization layer—either because it wasn't created with tf.get_variable() or your variable reuse logic is off. Here are actionable fixes to try:
1. Ensure All Variables Use tf.get_variable()
The error explicitly calls out that the missing variable might not have been created with tf.get_variable(). If you're using tf.Variable() anywhere in your BN layers or discriminator, swap it out for tf.get_variable(). This is critical because TensorFlow's variable scope reuse only works with variables created via this method.
For example, if you have a manual BN implementation:
# ❌ Bad: Uses tf.Variable which doesn't play nice with scope reuse beta = tf.Variable(tf.zeros([num_features]), name='beta') # ✅ Good: Uses tf.get_variable for scope compatibility beta = tf.get_variable('beta', shape=[num_features], initializer=tf.zeros_initializer())
If you're using tf.layers.batch_normalization, make sure you're not overriding the name parameter in a way that breaks scope tracking, and that the layer is created within a proper variable scope.
2. Fix Variable Scope Reuse Logic
Setting reuse=True globally is usually a mistake—you need to apply reuse only when you're reusing existing variables (like when running the discriminator on both real and fake images). Use tf.AUTO_REUSE instead of hardcoding True; it automatically reuses variables if they exist, or creates them if they don't.
Here's how to structure your discriminator calls correctly:
# First pass: Create discriminator variables for real images with tf.variable_scope('discriminator'): d_real = build_discriminator(real_images) # Second pass: Reuse existing variables for fake images with tf.variable_scope('discriminator', reuse=tf.AUTO_REUSE): d_fake = build_discriminator(fake_images)
Or, if your build_discriminator function takes a reuse parameter:
def build_discriminator(x, reuse=None): with tf.variable_scope('discriminator', reuse=reuse): # Your conv/BN/activation layers here return x # Create variables first d_real = build_discriminator(real_images) # Reuse variables for fake images d_fake = build_discriminator(fake_images, reuse=tf.AUTO_REUSE)
3. EMA Must Be Applied After Variable Creation
Your error triggers right before ema_apply_op, which suggests the EMA is trying to access variables that haven't been created yet. Always build your entire discriminator (and all its variables) before initializing the EMA and creating the apply operation.
Here's the correct order:
# Step 1: Build both discriminator passes to create all variables d_real = build_discriminator(real_images) d_fake = build_discriminator(fake_images, reuse=tf.AUTO_REUSE) # Step 2: Initialize EMA and target the discriminator's variables ema = tf.train.ExponentialMovingAverage(decay=0.999) # Get all trainable variables from the discriminator scope discriminator_vars = tf.get_collection(tf.GraphKeys.TRAINABLE_VARIABLES, scope='discriminator') # Create the EMA apply operation ema_apply_op = ema.apply(discriminator_vars)
This ensures the EMA is only trying to operate on variables that already exist in the graph.
4. Check for Scope Naming Conflicts
The error mentions d_bn1/d_bn1_2/moments/Squeeze/ExponentialMovingAverage—that nested naming suggests you might have duplicate scope names in your BN layers. Double-check your discriminator's BN layer scopes to make sure they're unique and properly nested. For example, don't name two different BN layers d_bn1; use d_bn1, d_bn2, etc., to avoid naming collisions.
Quick Example of a Working BN + EMA Setup
Here's a snippet of a discriminator with proper BN and EMA integration:
def build_discriminator(x, reuse=None): with tf.variable_scope('discriminator', reuse=reuse): # Conv + BN Layer 1 with tf.variable_scope('conv1'): x = tf.layers.conv2d(x, 64, 4, strides=2, padding='same') x = tf.layers.batch_normalization(x, training=True, name='bn1') x = tf.nn.leaky_relu(x) # Conv + BN Layer 2 with tf.variable_scope('conv2'): x = tf.layers.conv2d(x, 128, 4, strides=2, padding='same') x = tf.layers.batch_normalization(x, training=True, name='bn2') x = tf.nn.leaky_relu(x) # Final classification layer x = tf.layers.flatten(x) x = tf.layers.dense(x, 1) return x # Build discriminator for real and fake images real_imgs = tf.placeholder(tf.float32, [None, 64, 64, 3]) fake_imgs = generator(z) # Assume your generator is defined elsewhere d_real = build_discriminator(real_imgs) d_fake = build_discriminator(fake_imgs, reuse=tf.AUTO_REUSE) # Set up EMA ema = tf.train.ExponentialMovingAverage(0.99) disc_vars = tf.get_collection(tf.GraphKeys.TRAINABLE_VARIABLES, 'discriminator') ema_op = ema.apply(disc_vars) # Training loop (simplified) with tf.Session() as sess: sess.run(tf.global_variables_initializer()) for epoch in range(100): # Train discriminator _, d_loss_val = sess.run([d_train_op, d_loss], feed_dict={real_imgs: real_batch, z: noise_batch}) # Apply EMA after discriminator update sess.run(ema_op) # Train generator...
Give these steps a try—most likely the issue is a combination of variable creation method and scope reuse timing.
内容的提问来源于stack exchange,提问作者Divyat

