关于ReLU函数反向传播的疑问:线性函数与ReLU导数的困惑
Great question! Let's break this down step by step to clear up the confusion between linear functions and ReLU when it comes to backpropagation.
Recap: The Linear Function Problem
You nailed the core issue with linear functions like A = cx: their derivative with respect to x is a constant c. This means during gradient descent, the adjustment made during backpropagation is always identical, no matter how much the input x changes (delta(x)). Whether x is 1, 100, or -50, the gradient never shifts—so the weight update step size stays fixed. This is why linear functions can't learn complex, input-dependent patterns in neural networks.
How ReLU Differs
ReLU is defined as f(x) = max(0, x), and its derivative has two distinct, input-dependent cases:
- When
x > 0: The derivativef’(x) = 1(a constant, but only in this region) - When
x ≤ 0: The derivativef’(x) = 0
Here's the critical distinction from linear functions: ReLU's gradient isn't globally constant. It directly depends on the value of the input x:
- For positive inputs (
x > 0): The gradient is 1, so backprop will adjust weights using this constant—but only for inputs that "activate" the neuron. - For non-positive inputs (
x ≤ 0): The gradient drops to 0, meaning no weight updates occur for that neuron in this pass.
This means the backpropagation adjustment does depend on the input's value:
- Positive inputs trigger a consistent adjustment (similar to linear, but only in this subset)
- Non-positive inputs trigger no adjustment at all
Unlike linear functions, where every input leads to the same fixed update, ReLU's behavior shifts based on whether the input is activating the neuron. This input-dependent gradient is exactly why ReLU avoids the "one-size-fits-all" update problem of linear functions, and why it can learn non-linear patterns in data.
Quick Takeaway
- Linear functions have a global constant gradient, so backprop adjustments are identical for all inputs—this limits their ability to learn complexity.
- ReLU has a piecewise constant gradient that switches based on input value, so adjustments depend on whether the input is positive or not. This makes ReLU capable of modeling non-linear relationships, which is essential for most neural network tasks.
内容的提问来源于stack exchange,提问作者Hakan Utku Özdemir

