You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CNN全连接层后粗略深度图(Coarse7)生成逻辑技术问询

Hey there! Let's break down how the Coarse7 depth map is generated in that single-image depth prediction work, and fix up your Keras model to actually output that rough depth map instead of a 4096-dimensional feature vector.

First, Recap the Coarse Network Logic from the Paper

The Coarse7 branch is designed to output a low-resolution, global depth estimate. Here's the core idea:

  • It uses a modified AlexNet-style convolutional backbone to extract hierarchical image features.
  • After the final convolutional/pooling layers, the flattened feature vector feeds into fully connected layers that map directly to the flattened pixels of the coarse depth map.
  • The final step reshapes that 1D vector back into a 2D single-channel depth map—this is your Coarse7 output.

What's Missing in Your Current Model

Looking at your model summary, your final layer outputs a (None, 4096) tensor—this is just a feature vector, not a depth map. You're missing two critical pieces:

  1. The final dense layer needs to output exactly the number of pixels in your target coarse depth map (height × width, since it's a single-channel map).
  2. You need to reshape that 1D tensor into a 2D spatial image structure.

Step-by-Step Fix for Your Model

Let's tailor this to your input image size ((172, 576, 3)—I inferred this from your first Conv2D output shape (43, 144, 96) since stride=4 reduces each dimension by 4x).

1. Calculate Your Target Coarse Depth Map Size

Based on your existing pooling layers:

  • After Conv1 + Pool1: (21, 72)
  • After Conv2 + Pool2: (10, 36)
  • After the final MaxPooling2D: (5, 18)

For Coarse7, a common choice is 1/8 of the input size ((21, 72) here) since it balances detail and computational cost.

2. Modify Your Model Code

Add these layers at the end of your existing model to generate the Coarse7 depth map:

# Replace your final Dense(4096) with a layer that outputs the number of pixels in your target depth map
model.add(Dense(21 * 72))  # 21*72 = 1512 neurons for a 21x72 single-channel map
model.add(Activation('linear'))  # Use linear activation for unbounded depth values (relu clips to positive, which is okay for depth)
model.add(tf.keras.layers.Reshape((21, 72, 1)))  # Reshape the 1D vector into a 2D depth map

If you prefer a smaller coarse map (like the (5,18) size from your final pooling), adjust the Dense layer to 5*18=90 neurons and Reshape to (5,18,1).

3. Optional: Align with the Paper's Exact Structure

The original paper's Coarse network includes a 5th convolutional layer (Conv5) before pooling and fully connected layers. If you want to match that exactly, insert this after your Conv4 layer:

# Add Conv5 (matches original paper's structure)
model.add(Conv2D(256, (3,3), padding='same'))
model.add(Activation("relu"))
model.add(MaxPooling2D(pool_size=(2,2)))

Why This Works

  • The final dense layer translates the global convolutional features into individual depth values for every pixel in the coarse map.
  • The Reshape layer converts the flat list of depth values back into a spatial structure that represents the scene's rough depth layout. This Coarse7 map is then used as input to the Fine network to refine into a higher-resolution depth map.

内容的提问来源于stack exchange,提问作者jafar ali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:21:23