卷积神经网络中的反向传播及滤波器更新方法咨询
Hey there! Let's work through your CNN backpropagation confusion step by step—no fancy frameworks like TensorFlow, just plain, Python-focused explanations that tie everything together.
Filters are just weight matrices—but with two key rules that make CNNs efficient:
- Shared Weights: The same filter is reused across the entire input feature map. Unlike a regular NN where every input neuron has its own unique weight, a single filter's weights apply to every local patch (receptive field) of the input.
- Sparse Connectivity: Each output neuron only connects to a small local region of the input (the receptive field matching the filter size), not every input neuron. This cuts down on both parameters and computation.
They don't have a complex "network structure" of their own—they're just sets of weights designed to detect specific local patterns (like edges or textures) across the entire input.
You already get fully connected layer backprop, so let's start from there and work backwards to the convolutional layers.
1. Pass Error from the Fully Connected Layer to the Last Convolutional Layer
The final convolutional layer's output is usually flattened into a 1D vector to feed into the fully connected layer. To pass error back, take the error gradient from the fully connected layer (dL/dy, where y is the flattened input to the FC layer) and reshape it back into the original 3D shape of the convolutional layer's output (channels × height × width). This gives you dL/dZ, where Z is the convolutional layer's linear output (before activation).
2. Calculating Gradients for Filter Weights (Your Biggest Question!)
This is the heart of it. For each filter W, its gradient dL/dW is calculated using a cross-correlation between the input feature map and the error feature map dZ. Here's why:
- During forward propagation, each output value Z[i][j] is the dot product of the filter and the corresponding receptive field in the input.
- During backprop, by the chain rule, the gradient for each weight in the filter is the sum of (dL/dZ[i][j] × the input value that the weight multiplied with in forward prop) for every position [i][j] where that weight was used.
In plain terms: you slide the error feature map over the input feature map, multiply overlapping elements, and sum them up for each weight in the filter.
Simplified Python Pseudocode for Filter Gradients
import numpy as np def conv_backward_weights(input_feature_map, error_map, filter_shape, stride=1): out_channels, in_channels, filter_h, filter_w = filter_shape dW = np.zeros(filter_shape) for out_c in range(out_channels): for in_c in range(in_channels): for h in range(filter_h): for w in range(filter_w): # Sum the product of input patches and error map positions where this weight was used dW[out_c, in_c, h, w] = np.sum( input_feature_map[in_c, h::stride, w::stride] * error_map[out_c, :, :] ) return dW
3. Calculating Gradients for Biases
Bias gradients are straightforward: for each output channel, sum all the error values in the corresponding error map. Since every output value in the channel shares the same bias, all error contributions add up to a single value per channel.
def conv_backward_bias(error_map): # Sum over height and width for each output channel db = np.sum(error_map, axis=(1, 2)) return db
4. Passing Error to the Previous Convolutional Layer
If there's another convolutional layer before this one, you need to propagate the error backward to it. This involves:
- Rotating the filter 180 degrees (to reverse the forward convolution direction).
- Performing a convolution between the rotated filter and the current error map (this is often called "transpose convolution," but it's just error propagation, not actual deconvolution).
- Applying the derivative of the previous layer's activation function (e.g., for ReLU, zero out errors where the previous layer's linear output was negative).
Simplified Pseudocode for Error Propagation to Previous Layer
def conv_backward_input(error_map, filters, stride=1): out_channels, in_channels, filter_h, filter_w = filters.shape out_h, out_w = error_map.shape[1], error_map.shape[2] in_h = (out_h - 1) * stride + filter_h in_w = (out_w - 1) * stride + filter_w dX = np.zeros((in_channels, in_h, in_w)) # Rotate filters 180 degrees to reverse forward convolution rotated_filters = np.rot90(filters, 2, axes=(2, 3)) for in_c in range(in_channels): for out_c in range(out_channels): for h in range(out_h): for w in range(out_w): # Add error contribution to the corresponding receptive field in the input dX[in_c, h*stride:h*stride+filter_h, w*stride:w*stride+filter_w] += ( error_map[out_c, h, w] * rotated_filters[out_c, in_c] ) return dX
- Sparse Connectivity: Since each output neuron only connects to a local input patch, during backprop, each error value only affects weights in the corresponding filter's receptive field—you don't have to compute gradients for every input-output pair like in a fully connected layer.
- Shared Weights: The same filter is used across the entire input, so every position where the filter was applied contributes to the same weight's gradient. That's why we sum over all those positions in the filter gradient calculation—it's the shared weights rule in action.
Start small:
- Implement forward convolution with no padding, stride=1.
- Add a simple activation (like ReLU) and its derivative for backprop.
- Implement the filter and bias gradient functions above.
- Tie it to a small fully connected layer at the end, using the backprop you already know.
You don't need complex math—just stick to the chain rule and the intuition of how each weight contributes to the output during forward pass, then reverse that for backprop.
内容的提问来源于stack exchange,提问作者Fcoder

