You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CNN输出ndarray转图像异常问题排查求助

Hey there, let's break down your deraining CNN issue and walk through actionable fixes and troubleshooting steps:

Key Troubleshooting Directions & Fixes

1. Root Cause: Model Output Compression from Unnecessary Max Pool

The most likely culprit is the max_pool layer added after your transposed convolution in conv_6. Here's why:

  • Transposed convolution is meant to restore the original image size and spatial details. Adding a max pool here (with stride=1, padding=SAME) will only retain the maximum value in each 3x3 window, crushing most pixel values to near-zero—this directly causes the dark output you're seeing.

Fix: Remove the Final Max Pool & Add Sigmoid Activation

Replace your conv_6 output code with this:

with tf.variable_scope('conv_6'):
    kernel = tf.Variable(tf.random_normal([8, 8, FLAGS.num_channels,FLAGS.num_feature], dtype=tf.float32, stddev=1e-3), trainable=True, name='weights3')
    biases = tf.Variable(tf.constant(0.0, shape=[FLAGS.num_channels], dtype=tf.float32), trainable=True, name='biases3')
    conv = tf.nn.conv2d_transpose(conv4,kernel,[tf.shape(images)[0], tf.shape(images)[1], tf.shape(images)[2], FLAGS.num_channels],[1, 1, 1, 1],padding='SAME')
    conv5 = tf.nn.bias_add(conv, biases)
    out = tf.nn.sigmoid(conv5)  # Replace max_pool with sigmoid to clamp values to [0,1]
  • Sigmoid ensures your output stays strictly within the [0,1] range (perfect for image data) without destroying spatial details.

2. Activation Function Choice is Hurting Output Range

You're using tanh for all intermediate layers, which outputs values in [-1,1]. Stacking multiple tanh layers can cause:

  • Gradient vanishing during training, making it hard for the model to learn proper brightness levels
  • Output values clustered near 0 even after clamping, leading to dark images

Fix: Swap Tanh with ReLU/LeakyReLU

Update your intermediate conv layers to use ReLU (or LeakyReLU for better gradient flow):

# Replace tanh with relu in conv1-conv4
conv1 = tf.nn.relu(bias)
# Or use LeakyReLU if you need to preserve negative gradients:
conv1 = tf.nn.leaky_relu(bias, alpha=0.2)

3. Debug Output Value Distribution

Before saving the image, add these lines to confirm if your output is actually in a visible range:

print(f"Derained stats - Min: {derained.min()}, Max: {derained.max()}, Mean: {derained.mean()}")
  • If the mean is close to 0, your model is indeed outputting mostly dark pixels—this confirms the issue is with the model, not the image saving code.

4. Fix Image Saving & Display Issues

Your current img.imsave adds an alpha channel (4th channel) when saving as PNG. To avoid this and ensure correct color space handling:

Option 1: Use OpenCV for Saving

# Convert [0,1] float to [0,255] uint8, and fix RGB->BGR for OpenCV
derained_uint8 = (derained * 255).astype(np.uint8)
cv2.imwrite("h_11.jpg", cv2.cvtColor(derained_uint8, cv2.COLOR_RGB2BGR))

Option 2: Force JPG Format with Matplotlib

img.imsave("h_11.jpg", derained, format='jpg')  # JPG doesn't support alpha channels

5. Double-Check Transposed Convolution Parameters

Your transposed convolution kernel shape [8,8,3,num_feature] is correct (matches TensorFlow's [height, width, output_channels, input_channels] requirement), but an 8x8 kernel might be overly large and cause excessive smoothing. Try reducing it to 3x3 or 4x4 for finer detail retention.


内容的提问来源于stack exchange,提问作者致翔黃

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:53:13