Caffe中TanHLayer的Backward_cpu为何针对sigmoid输入计算梯度?
Hey there! Let me clear up this confusion for you as someone new to Caffe. The reason the Backward_cpu() method for TanHLayer is described as calculating the error gradient with respect to the sigmoid input instead of the tanh input boils down to code reuse and the mathematical relationship between tanh and sigmoid functions.
Here's a breakdown:
First, remember the mathematical link between tanh and sigmoid:
tanh(x) = 2 * sigmoid(2x) - 1
This means tanh can be computed using the sigmoid function with a simple input transformation. Instead of writing entirely new forward and backward pass code for tanh, Caffe reuses the existingSigmoidLayerlogic to avoid redundancy.In the forward pass of
TanHLayer, the inputxis first transformed toz = 2x, then fed into the sigmoid function to getsigmoid(z), and finally converted totanh(x)using the formula above.For the backward pass, the chain rule comes into play. Suppose we have the upstream gradient
dL/dy(wherey = tanh(x)). To getdL/dx(the gradient with respect to the original tanh input), we can leverage the sigmoid's backward pass:- First compute
dL/dz(the gradient with respect to the sigmoid inputz), which is what the documentation refers to. This uses the sigmoid gradient formula:dL/dz = dL/dy * 2 * sigmoid(z) * (1 - sigmoid(z)). - Then convert
dL/dztodL/dxby multiplying by2(sincez = 2x, sodz/dx = 2).
- First compute
The documentation's description focuses on this internal step of calculating dL/dz, but the end result of Backward_cpu() is still the gradient with respect to the original tanh input x—it just uses the sigmoid's gradient calculation under the hood.
A great way to confirm this is to look at the source code for TanHLayer (check caffe/src/caffe/layers/tanh_layer.cpp). You'll see the backward pass explicitly handles this transformation between the sigmoid input gradient and the tanh input gradient.
内容的提问来源于stack exchange,提问作者ykmr

