You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ReLU激活网络反向传播中权重梯度为对角矩阵的直观解析与疑问

Understanding Diagonal Gradient Matrices in ReLU Layers

Great question—this is such a common "wait, why does that happen?" moment when diving into backprop for ReLU networks. Let’s unpack this step by step to build intuition:

First, let’s clarify the core misconception

A diagonal gradient matrix does NOT mean non-diagonal weights have no impact on loss overall. It only means that, for the current input batch/instance, those weights have a zero instantaneous gradient—so gradient descent won’t update them right now.

Why does the gradient matrix become diagonal?

This ties directly to ReLU’s sparse activation behavior. Let’s break this down in plain terms:

  • For a fully connected layer followed by ReLU, we have z = Wx + b (weighted sum of inputs) and a = max(0, z) (ReLU output).
  • During backpropagation, the gradient of loss with respect to z is ∇zL = ∇aL * I(z > 0), where I is an indicator function (1 if z > 0, 0 otherwise). This means neurons that are "off" (z ≤ 0) don’t pass any gradient backward to the weights that feed into them.
  • The gradient of loss with respect to weights W is calculated as ∇WL = ∇zL * x^T.

This becomes a diagonal matrix when:

  1. For each active output neuron (z > 0), only the corresponding input feature is non-zero (like one-hot encoded inputs), OR
  2. Each active neuron only receives a non-zero signal from a single input feature in the current batch.

In these scenarios, the only non-zero entries in ∇WL are where the row (output neuron index) matches the column (input feature index)—resulting in a diagonal matrix.

Do non-diagonal weights never get updated?

Absolutely not! The diagonal gradient is just a snapshot for the current input. If you feed a different input where:

  • A non-diagonal weight contributes to an active neuron (i.e., the weighted sum for that neuron is positive), AND
  • The input feature corresponding to that weight’s column is non-zero,

then that weight will get a non-zero gradient and be updated by gradient descent.

The only edge case where a weight might never update is if it’s part of a "dead ReLU"—a neuron that never activates for any training input (usually due to poor initialization that makes z always negative). But this is avoidable with proper initialization (like He initialization) or using variants like leaky ReLU.

Quick concrete example

Suppose we have a 2x2 weight matrix W, and input x = [1, 0]. If W[0][0] > 0 (first neuron activates) and W[1][0] < 0 (second neuron is off), then ∇WL will be:

[ [∇a0L * 1, 0],
  [0, 0] ]

Only the (0,0) diagonal entry is non-zero. But if we next feed x = [0, 1] and W[0][1] > 0, the gradient matrix becomes:

[ [0, ∇a0L * 1],
  [0, 0] ]

Now the non-diagonal (0,1) entry is non-zero, so W[0][1] gets updated.

内容的提问来源于stack exchange,提问作者djpetti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:46:51