CNN输出ndarray转图像异常问题排查求助
Hey there, let's break down your deraining CNN issue and walk through actionable fixes and troubleshooting steps:
1. Root Cause: Model Output Compression from Unnecessary Max Pool
The most likely culprit is the max_pool layer added after your transposed convolution in conv_6. Here's why:
- Transposed convolution is meant to restore the original image size and spatial details. Adding a max pool here (with stride=1, padding=SAME) will only retain the maximum value in each 3x3 window, crushing most pixel values to near-zero—this directly causes the dark output you're seeing.
Fix: Remove the Final Max Pool & Add Sigmoid Activation
Replace your conv_6 output code with this:
with tf.variable_scope('conv_6'): kernel = tf.Variable(tf.random_normal([8, 8, FLAGS.num_channels,FLAGS.num_feature], dtype=tf.float32, stddev=1e-3), trainable=True, name='weights3') biases = tf.Variable(tf.constant(0.0, shape=[FLAGS.num_channels], dtype=tf.float32), trainable=True, name='biases3') conv = tf.nn.conv2d_transpose(conv4,kernel,[tf.shape(images)[0], tf.shape(images)[1], tf.shape(images)[2], FLAGS.num_channels],[1, 1, 1, 1],padding='SAME') conv5 = tf.nn.bias_add(conv, biases) out = tf.nn.sigmoid(conv5) # Replace max_pool with sigmoid to clamp values to [0,1]
- Sigmoid ensures your output stays strictly within the [0,1] range (perfect for image data) without destroying spatial details.
2. Activation Function Choice is Hurting Output Range
You're using tanh for all intermediate layers, which outputs values in [-1,1]. Stacking multiple tanh layers can cause:
- Gradient vanishing during training, making it hard for the model to learn proper brightness levels
- Output values clustered near 0 even after clamping, leading to dark images
Fix: Swap Tanh with ReLU/LeakyReLU
Update your intermediate conv layers to use ReLU (or LeakyReLU for better gradient flow):
# Replace tanh with relu in conv1-conv4 conv1 = tf.nn.relu(bias) # Or use LeakyReLU if you need to preserve negative gradients: conv1 = tf.nn.leaky_relu(bias, alpha=0.2)
3. Debug Output Value Distribution
Before saving the image, add these lines to confirm if your output is actually in a visible range:
print(f"Derained stats - Min: {derained.min()}, Max: {derained.max()}, Mean: {derained.mean()}")
- If the mean is close to 0, your model is indeed outputting mostly dark pixels—this confirms the issue is with the model, not the image saving code.
4. Fix Image Saving & Display Issues
Your current img.imsave adds an alpha channel (4th channel) when saving as PNG. To avoid this and ensure correct color space handling:
Option 1: Use OpenCV for Saving
# Convert [0,1] float to [0,255] uint8, and fix RGB->BGR for OpenCV derained_uint8 = (derained * 255).astype(np.uint8) cv2.imwrite("h_11.jpg", cv2.cvtColor(derained_uint8, cv2.COLOR_RGB2BGR))
Option 2: Force JPG Format with Matplotlib
img.imsave("h_11.jpg", derained, format='jpg') # JPG doesn't support alpha channels
5. Double-Check Transposed Convolution Parameters
Your transposed convolution kernel shape [8,8,3,num_feature] is correct (matches TensorFlow's [height, width, output_channels, input_channels] requirement), but an 8x8 kernel might be overly large and cause excessive smoothing. Try reducing it to 3x3 or 4x4 for finer detail retention.
内容的提问来源于stack exchange,提问作者致翔黃

