为何WaveNet与PixelCNN中门控激活函数性能优于ReLU?
Great question—this is a common point of confusion when working through WaveNet and PixelCNN papers. Let’s unpack why that gated activation outperforms ReLU, and why it’s definitely not just a lucky accident.
First, let’s clarify the gated activation function you referenced (using ⊙ for element-wise multiplication to make the operation clearer):
y = tanh(W_{k,f} * x) ⊙ σ(W_{k,g} * x)
Here, W_{k,f} and W_{k,g} are convolution kernels for the "filter" and "gate" branches, tanh squashes input values to the range [-1,1], and σ (sigmoid) outputs a gate value between 0 and 1 to control how much of the tanh output is passed through.
Key Reasons It Outperforms ReLU
Adaptive, fine-grained control over information flow
ReLU uses a rigid hard threshold: any input value below 0 gets zeroed out, with no room for nuance. The gated activation is dynamic: the sigmoid gate learns to selectively pass information based on the input context. For tasks like audio generation (WaveNet) or pixel-by-pixel image generation (PixelCNN), this is critical—different parts of the input sequence or pixel grid have varying importance, and the gate lets the model tune which signals to keep or discard.Better gradient preservation without sacrificing non-linearity
ReLU suffers from "dying neurons": if a neuron’s input is consistently negative, its gradient becomes zero, and it stops learning entirely. Tanh alone has gradient vanishing issues in its saturated regions, but pairing it with a sigmoid gate mitigates this problem. When the gate is close to 1, the gradient of the tanh branch is preserved; when the gate is close to 0, that branch is effectively turned off, so we don’t waste computation or propagate noisy gradients. Plus, the combination of tanh and sigmoid delivers a richer non-linear transformation than ReLU’s simple piecewise linearity.Aligned with the core needs of sequence/generative tasks
Gated mechanisms were already proven effective in recurrent models like LSTMs and GRUs, where they handle long-term dependencies by controlling memory retention. WaveNet and PixelCNN adapted this idea to convolutional models: instead of managing temporal memory, they control spatial/temporal information flow via convolutions. This isn’t a random experiment—it’s a deliberate extension of a known successful pattern to a new model architecture.
Is it just an accidental improvement?
Absolutely not. Researchers already knew that gated mechanisms excel at tasks where information needs to be selectively passed or retained across sequences (or spatial grids, in PixelCNN’s case). Applying this concept to convolutional layers was a logical next step, not a random swap of activation functions. The empirical performance boost confirmed that the gated approach directly addresses key limitations of ReLU in these specific generative tasks.
内容的提问来源于stack exchange,提问作者Ahmad Moussa

