You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow下CNN技术问题:获取隐藏FC层编码与自动计算FC输入尺寸

Answers to Your TensorFlow CNN Questions

Hey Francis, great job getting reasonable results with your custom LeNet-5 on your image dataset—let's clean up those rough edges with solutions to your two questions:


1. Extract the Last Hidden FC Layer as Data Encoding

Your last hidden fully connected layer is layer4_actv (since logits is the final classification output). To use this as an encoding for other classifiers, you just need to modify your model_lenet5 function to return both the logits and this layer's output. Here's the updated code:

def model_lenet5(data, variables):
    layer1_conv = tf.nn.conv2d(data, variables['w1'], [1, 1, 1, 1], padding='SAME')
    layer1_actv = tf.sigmoid(layer1_conv + variables['b1'])
    layer1_pool = tf.nn.avg_pool(layer1_actv, [1, 2, 2, 1], [1, 2, 2, 1], padding='SAME')
    
    layer2_conv = tf.nn.conv2d(layer1_pool, variables['w2'], [1, 2, 2, 1], padding='SAME')
    layer2_actv = tf.sigmoid(layer2_conv + variables['b2'])
    layer2_pool = tf.nn.max_pool(layer2_actv, [1, 2, 2, 1], [1, 2, 2, 1], padding='SAME')
    
    flat_layer = flatten_tf_array(layer2_pool)
    layer3_fccd = tf.matmul(flat_layer, variables['w3']) + variables['b3']
    layer3_actv = tf.nn.sigmoid(layer3_fccd)
    
    layer4_fccd = tf.matmul(layer3_actv, variables['w4']) + variables['b4']
    layer4_actv = tf.nn.sigmoid(layer4_fccd)  # This is your encoding layer
    
    logits = tf.matmul(layer4_actv, variables['w5']) + variables['b5']
    
    # Return both logits (for training) and the encoding (for other classifiers)
    return logits, layer4_actv

How to Use It:

When you run your model (either training or inference), capture both outputs:

# Example during inference
logits, encoding = model_lenet5(test_data, vars)
# Now you can pass `encoding` to scikit-learn classifiers like SVM, Random Forest, etc.

2. Automatically Calculate Input Size for the First FC Layer

Hardcoding 288 in w3 is error-prone when changing image sizes or network architecture. Instead, calculate the flattened size programmatically based on your input image dimensions and convolution/pooling parameters. Here's how to update your variables_lenet5 function:

def variables_lenet5(filter_size=filter_size_1, filter_size_2=filter_size_2, 
                     filter_depth1=filter_size_1, filter_depth2=filter_size_2, 
                     num_hidden1=hid_1, num_hidden2=hid_2, 
                     image_width=image_width, image_height=image_height, 
                     image_depth=3, num_labels=num_labels):
    # Calculate dimensions after each convolution/pooling step
    # Step 1: Conv1 + AvgPool
    # Conv1 uses SAME padding, stride 1 → size stays same as input
    pool1_width = (image_width + 1) // 2  # AvgPool 2x2, stride 2, SAME padding
    pool1_height = (image_height + 1) // 2
    
    # Step 2: Conv2 + MaxPool
    # Conv2 uses SAME padding, stride 2 → size is halved
    conv2_width = (pool1_width + 1) // 2
    conv2_height = (pool1_height + 1) // 2
    # MaxPool 2x2, stride 2, SAME padding → size halved again
    pool2_width = (conv2_width + 1) // 2
    pool2_height = (conv2_height + 1) // 2
    
    # Flattened size = height * width * number of filters from last conv layer
    flatten_size = pool2_width * pool2_height * filter_depth2
    
    # Now create variables with the calculated flatten_size
    w1 = tf.Variable(tf.truncated_normal([filter_size, filter_size, image_depth, filter_depth1], stddev=0.1))
    b1 = tf.Variable(tf.zeros([filter_depth1]))
    
    w2 = tf.Variable(tf.truncated_normal([filter_size_2, filter_size_2, filter_depth1, filter_depth2], stddev=0.1))
    b2 = tf.Variable(tf.constant(1.0, shape=[filter_depth2]))
    
    # Use flatten_size instead of hardcoded 288
    w3 = tf.Variable(tf.truncated_normal([flatten_size, num_hidden1], stddev=0.1))
    b3 = tf.Variable(tf.constant(1.0, shape=[num_hidden1]))
    
    w4 = tf.Variable(tf.truncated_normal([num_hidden1, num_hidden2], stddev=0.1))
    b4 = tf.Variable(tf.constant(1.0, shape=[num_hidden2]))
    
    w5 = tf.Variable(tf.truncated_normal([num_hidden2, num_labels], stddev=0.1))
    b5 = tf.Variable(tf.constant(1.0, shape=[num_labels]))
    
    variables = {
        'w1': w1, 'w2': w2, 'w3': w3, 'w4': w4, 'w5': w5,
        'b1': b1, 'b2': b2, 'b3': b3, 'b4': b4, 'b5': b5
    }
    return variables

Key Notes:

  • The (size + 1) // 2 calculation handles both even and odd dimensions correctly for SAME padding with stride 2.
  • If you ever modify your convolution/pooling strides or padding, just update these dimension calculations—no more manual arithmetic!

内容的提问来源于stack exchange,提问作者FrancisSem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:02:01