You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于卷积层反向传播实现方法的技术问询

Great question—you’re absolutely right that while forward-pass convolution gets tons of visual love in tutorials, the backward pass is often overlooked, especially in intuitive, visual explanations. Let’s break this down step by step, starting with your solid understanding of fully connected layers to draw clear parallels.

First, Confirming Your Fully Connected Intuition

You nailed the chain rule for fully connected layers. For a layer where ( H_1 = W_1 \cdot I_1 + b_1 ) (simplified to a single input ( I_1 )), the gradient of the error with respect to the weight ( W_1 ) is:
$$ \frac{\partial Error}{\partial W_1} = \frac{\partial Error}{\partial HA_1} \cdot \frac{\partial HA_1}{\partial H_1} \cdot \frac{\partial H_1}{\partial W_1} $$
And as you noted, ( \frac{\partial H_1}{\partial W_1} = I_1 )—the input value paired with that weight. This core idea (using the layer’s input to compute weight gradients) translates directly to convolution, but we have to adjust for the sliding window nature of conv layers.


Convolutional Layer Backward Pass: Step by Step

Let’s use a simple 2D example to make this concrete:

  • Input feature map ( I ): 3x3 (values ( I_{i,j} ))
  • Conv kernel ( K ): 2x2 (values ( K_{m,n} ))
  • Output feature map ( O ): 2x2 (no padding, stride=1)
  • Forward pass: ( O_{x,y} = \sum_{m=0}^1 \sum_{n=0}^1 I_{x+m, y+n} \cdot K_{m,n} )

1. Gradient of Error w.r.t. Kernel (( \frac{\partial Error}{\partial K} ))

This is the equivalent of ( \frac{\partial Error}{\partial W_1} ) in fully connected layers. For each kernel element ( K_{m,n} ), its gradient is the sum of products between:

  • Every input pixel ( I_{x+m, y+n} ) that ( K_{m,n} ) multiplied during the forward pass
  • The corresponding error gradient from the output ( \frac{\partial Error}{\partial O_{x,y}} )

Mathematically:
$$ \frac{\partial Error}{\partial K_{m,n}} = \sum_{x=0}^1 \sum_{y=0}^1 \frac{\partial Error}{\partial O_{x,y}} \cdot I_{x+m, y+n} $$

Intuitive parallel: In fully connected layers, each weight only pairs with one input. In convolution, each kernel element pairs with multiple input pixels (one per sliding window), so we sum all those input-error products. Visually, this is like performing a valid cross-correlation between the input feature map and the error gradient map from the next layer.

2. Gradient of Error w.r.t. Input (( \frac{\partial Error}{\partial I} ))

If we need to propagate gradients back to the previous layer (e.g., from a conv layer to an input image or another conv layer), we compute this by "spreading" the output error gradients using the kernel—specifically, a full convolution with the flipped kernel.

For an input pixel ( I_{i,j} ), its gradient is the sum of products between:

  • Every kernel element ( K_{i-x, j-y} ) that multiplied ( I_{i,j} ) during the forward pass (this is the flipped kernel)
  • The corresponding output error gradient ( \frac{\partial Error}{\partial O_{x,y}} )

Mathematically:
$$ \frac{\partial Error}{\partial I_{i,j}} = \sum_{x,y} \frac{\partial Error}{\partial O_{x,y}} \cdot K_{i-x, j-y} $$
(We only include terms where ( x,y ) are valid indices for ( O ), i.e., where the kernel window included ( I_{i,j} ) in the forward pass.)

Intuitive parallel: This is like how in fully connected layers, you compute input gradients by multiplying weight values with downstream error gradients—here, we just account for the kernel sliding over multiple positions.

3. Gradient of Error w.r.t. Bias (if used)

This is straightforward, just like in fully connected layers: the bias is added to every output element, so its gradient is the sum of all output error gradients:
$$ \frac{\partial Error}{\partial b} = \sum_{x,y} \frac{\partial Error}{\partial O_{x,y}} $$


Key Takeaways

  • Convolutional backward pass leans on the same chain rule logic as fully connected layers, but adapts to the kernel’s sliding window behavior.
  • Kernel gradients = sum of input-error products across all windows where the kernel element was used.
  • Input gradients = spread error gradients using the flipped kernel (to reverse the forward pass’s sliding).

内容的提问来源于stack exchange,提问作者Edv Beq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:23:37