TensorFlow下CNN技术问题:获取隐藏FC层编码与自动计算FC输入尺寸
Hey Francis, great job getting reasonable results with your custom LeNet-5 on your image dataset—let's clean up those rough edges with solutions to your two questions:
1. Extract the Last Hidden FC Layer as Data Encoding
Your last hidden fully connected layer is layer4_actv (since logits is the final classification output). To use this as an encoding for other classifiers, you just need to modify your model_lenet5 function to return both the logits and this layer's output. Here's the updated code:
def model_lenet5(data, variables): layer1_conv = tf.nn.conv2d(data, variables['w1'], [1, 1, 1, 1], padding='SAME') layer1_actv = tf.sigmoid(layer1_conv + variables['b1']) layer1_pool = tf.nn.avg_pool(layer1_actv, [1, 2, 2, 1], [1, 2, 2, 1], padding='SAME') layer2_conv = tf.nn.conv2d(layer1_pool, variables['w2'], [1, 2, 2, 1], padding='SAME') layer2_actv = tf.sigmoid(layer2_conv + variables['b2']) layer2_pool = tf.nn.max_pool(layer2_actv, [1, 2, 2, 1], [1, 2, 2, 1], padding='SAME') flat_layer = flatten_tf_array(layer2_pool) layer3_fccd = tf.matmul(flat_layer, variables['w3']) + variables['b3'] layer3_actv = tf.nn.sigmoid(layer3_fccd) layer4_fccd = tf.matmul(layer3_actv, variables['w4']) + variables['b4'] layer4_actv = tf.nn.sigmoid(layer4_fccd) # This is your encoding layer logits = tf.matmul(layer4_actv, variables['w5']) + variables['b5'] # Return both logits (for training) and the encoding (for other classifiers) return logits, layer4_actv
How to Use It:
When you run your model (either training or inference), capture both outputs:
# Example during inference logits, encoding = model_lenet5(test_data, vars) # Now you can pass `encoding` to scikit-learn classifiers like SVM, Random Forest, etc.
2. Automatically Calculate Input Size for the First FC Layer
Hardcoding 288 in w3 is error-prone when changing image sizes or network architecture. Instead, calculate the flattened size programmatically based on your input image dimensions and convolution/pooling parameters. Here's how to update your variables_lenet5 function:
def variables_lenet5(filter_size=filter_size_1, filter_size_2=filter_size_2, filter_depth1=filter_size_1, filter_depth2=filter_size_2, num_hidden1=hid_1, num_hidden2=hid_2, image_width=image_width, image_height=image_height, image_depth=3, num_labels=num_labels): # Calculate dimensions after each convolution/pooling step # Step 1: Conv1 + AvgPool # Conv1 uses SAME padding, stride 1 → size stays same as input pool1_width = (image_width + 1) // 2 # AvgPool 2x2, stride 2, SAME padding pool1_height = (image_height + 1) // 2 # Step 2: Conv2 + MaxPool # Conv2 uses SAME padding, stride 2 → size is halved conv2_width = (pool1_width + 1) // 2 conv2_height = (pool1_height + 1) // 2 # MaxPool 2x2, stride 2, SAME padding → size halved again pool2_width = (conv2_width + 1) // 2 pool2_height = (conv2_height + 1) // 2 # Flattened size = height * width * number of filters from last conv layer flatten_size = pool2_width * pool2_height * filter_depth2 # Now create variables with the calculated flatten_size w1 = tf.Variable(tf.truncated_normal([filter_size, filter_size, image_depth, filter_depth1], stddev=0.1)) b1 = tf.Variable(tf.zeros([filter_depth1])) w2 = tf.Variable(tf.truncated_normal([filter_size_2, filter_size_2, filter_depth1, filter_depth2], stddev=0.1)) b2 = tf.Variable(tf.constant(1.0, shape=[filter_depth2])) # Use flatten_size instead of hardcoded 288 w3 = tf.Variable(tf.truncated_normal([flatten_size, num_hidden1], stddev=0.1)) b3 = tf.Variable(tf.constant(1.0, shape=[num_hidden1])) w4 = tf.Variable(tf.truncated_normal([num_hidden1, num_hidden2], stddev=0.1)) b4 = tf.Variable(tf.constant(1.0, shape=[num_hidden2])) w5 = tf.Variable(tf.truncated_normal([num_hidden2, num_labels], stddev=0.1)) b5 = tf.Variable(tf.constant(1.0, shape=[num_labels])) variables = { 'w1': w1, 'w2': w2, 'w3': w3, 'w4': w4, 'w5': w5, 'b1': b1, 'b2': b2, 'b3': b3, 'b4': b4, 'b5': b5 } return variables
Key Notes:
- The
(size + 1) // 2calculation handles both even and odd dimensions correctly for SAME padding with stride 2. - If you ever modify your convolution/pooling strides or padding, just update these dimension calculations—no more manual arithmetic!
内容的提问来源于stack exchange,提问作者FrancisSem

