You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Gradient Tape计算LRP反向梯度全为0的问题求助

问题描述

我正尝试将LRP(Layer-Wise Relevance Propagation)输出的热力图反向映射到模型权重,计划通过最小化未训练模型权重的相关性与目标热力图相关性得分的损失,借助Gradient Tape让权重达到能生成目标热力图的取值。

基于MNIST数据集搭建的全连接模型:

num_classes = 10
input_layer = Input(shape=(img_width * img_height,))
X = Dense(300, activation='relu', kernel_regularizer='l2')(input_layer)
X = Dense(100, activation='relu',kernel_regularizer='l2')(X)
X = Dense(10, activation='relu',kernel_regularizer='l2')(X)

model = Model(inputs=input_layer, outputs=X)

实现的LRP函数:

def get_relevance_tf(W,B,img,pred):
    L = len(W)
    A = [img]+[None]*L
    for l in range(L):
        A[l+1] = tf.nn.relu(tf.matmul(A[l],W[l])+B[l])
    R = [0.0]*L + [A[L]*(pred)]
    for l in range(1,L)[::-1]:

        w = W[l]
        b = B[l]

        z = tf.matmul(A[l],w)+b    # step 1

        s = R[l+1] / z               # step 2
        c = tf.matmul(s,w,transpose_b=True)          # step 3
        R[l] = A[l]*c                # step 4
       
    w  = W[0]
    wp = tf.math.maximum(0,w)
    wm = tf.math.minimum(0,w)
    lb = A[0]*0-1
    hb = A[0]*0+1

    z = tf.matmul(A[0],w)-tf.matmul(lb,wp)-tf.matmul(hb,wm)+1e-9        # step 1
    s = R[1]/z                                        # step 2
    c,cp,cm  = tf.matmul(s,w,transpose_b=True),tf.matmul(s,wp,transpose_b=True),tf.matmul(s,wm,transpose_b=True) # step 3
    R[0] = A[0]*c-lb*cp-hb*cm                         # step 4
    return R

梯度计算代码:

model = tf.keras.models.load_model("model.10.hdf5")

img = tf.convert_to_tensor(X_train[index].reshape(1,784),dtype=tf.float32)
pred = tf.convert_to_tensor(y_train_one_hot[index],dtype=tf.float32)
W = [tf.Variable(i,dtype=tf.float32,trainable=True) for i in model.get_weights()[::2]]
B = [tf.Variable(i,dtype=tf.float32,trainable=True) for i in model.get_weights()[1::2]]
with tf.GradientTape() as tape:
    R = get_relevance_tf(W,B,img,pred)
    loss = tf.math.reduce_sum(tf.math.abs(R[0]-pred_R[0]))
    
grads = tape.gradient(loss, [W,B])
print(grads)

运行后所有权重梯度均为0:

[[<tf.Tensor: shape=(784, 300), dtype=float32, numpy=
array([[0., 0., 0., ..., 0., 0., 0.],
       [0., 0., 0., ..., 0., 0., 0.],
       [0., 0., 0., ..., 0., 0., 0.],
       ...,
       [0., 0., 0., ..., 0., 0., 0.],
       [0., 0., 0., ..., 0., 0., 0.],
       [0., 0., 0., ..., 0., 0., 0.]], dtype=float32)>, <tf.Tensor: shape=(300, 100), dtype=float32, numpy=
array([[0., 0., 0., ..., 0., 0., 0.],
       [0., 0., 0., ..., 0., 0., 0.],
       [0., 0., 0., ..., 0., 0., 0.],
       ...,
       [0., 0., 0., ..., 0., 0., 0.],
       [0., 0., 0., ..., 0., 0., 0.],
       [0., 0., 0., ..., 0., 0., 0.]], dtype=float32)>, <tf.Tensor: shape=(100, 10), dtype=float32, numpy=
array([[0., 0., 0., 0., 0., 0., 0., 0., 0., 0.],
       [0., 0., 0., 0., 0., 0., 0., 0., 0., 0.],
       [0., 0., 0., 0., 0., 0., 0., 0., 0., 0.],
       [0., 0., 0., 0., 0., 0., 0., 0., 0., 0.],
       [0., 0., 0., 0., 0., 0., 0., 0., 0., 0.],
       [0., 0., 0., 0., 0., 0., 0., 0., 0., 0.],
       [0., 0., 0., 0., 0., 0., 0., 0., 0., 0.],
       [0., 0., 0., 0., 0., 0., 0., 0., 0., 0.],
       [0., 0., 0., 0., 0., 0., 0., 0., 0., 0.],
       [0., 0., 0., 0., 0., 0., 0., 0., 0., 0.],
...
       0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.,
       0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.,
       0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.],
      dtype=float32)>, <tf.Tensor: shape=(10,), dtype=float32, numpy=array([-0., -0., -0., -0., -0., -0., -0., -0., -0., -0.], dtype=float32)>]]

已尝试autograd、自定义模型+自定义损失函数等方式,结果均相同,无法理解为何权重与相关性存在直接关联但梯度全为0。

问题分析与解决方案

梯度全为0的核心原因在于LRP计算流程中存在的不可导点或梯度消失场景,具体拆解如下:

1. ReLU激活函数的梯度截断

模型和LRP前向计算中都使用了tf.nn.relu,当神经元输出为0时,ReLU的梯度为0。如果LRP前向传播的某一层中,A[l+1] = tf.nn.relu(...)的输出大量为0,后续的相关性传播会因为这些0值导致梯度无法反向传递到权重。

解决方案:

  • 替换ReLU为带泄漏的LeakyReLU,保留小梯度避免截断:
    A[l+1] = tf.nn.leaky_relu(tf.matmul(A[l],W[l])+B[l], alpha=0.01)
    
  • 或者在LRP计算中,对ReLU的输出添加极小的偏移量(如1e-9),避免完全为0:
    A[l+1] = tf.nn.relu(tf.matmul(A[l],W[l])+B[l]) + 1e-9
    

2. LRP中的除法操作导致梯度消失

在LRP的反向传播步骤中,s = R[l+1] / z 当z的绝对值很大时,s会趋近于0,后续的梯度相乘后会被稀释为0;另外当z接近0时,虽然添加了1e-9的偏移,但数值不稳定可能导致梯度被强制置0。

解决方案:

  • 对z添加更稳定的数值偏移,或使用平滑除法替代普通除法:
    s = R[l+1] / (z + tf.sign(z)*1e-9)
    
  • 改用LRP-ε变体,在z中添加小正项避免除法不稳定:
    z = tf.matmul(A[l],w)+b + 1e-6  # 添加ε项
    

3. 权重拆分导致的梯度传递中断

在第一层LRP计算中,将权重拆分为wp(正权重)和wm(负权重)并分别计算cp和cm,如果某类权重的梯度在计算中被抵消,会导致最终权重的梯度为0。

解决方案:

  • 检查wp和wm的分布,若存在大量0或接近0的权重,初始化时调整权重分布避免极端值;
  • 改用更简洁的LRP-0规则计算第一层相关性,减少梯度传递分支:
    # 替换第一层的LRP计算为LRP-0规则
    z = tf.matmul(A[0], w) + b + 1e-9
    s = R[1] / z
    c = tf.matmul(s, w, transpose_b=True)
    R[0] = A[0] * c
    

4. 梯度磁带的追踪范围问题

创建的W和B是tf.Variable,需确保在GradientTape上下文内所有涉及这些变量的计算都被正确追踪,可显式设置强制追踪:

with tf.GradientTape() as tape:
    tape.watch(W + B)  # 显式追踪所有权重和偏置变量
    R = get_relevance_tf(W,B,img,pred)
    loss = tf.math.reduce_sum(tf.math.abs(R[0]-pred_R[0]))

验证步骤

  1. 简化模型,去掉正则化和部分层,用小全连接层测试梯度是否正常;
  2. 打印LRP计算中的z、s、c等中间值,查看是否存在大量0或极端值;
  3. 替换损失函数为MSE(tf.math.reduce_mean(tf.square(R[0]-pred_R[0]))),L1损失的梯度是阶跃函数,可能在某些点导致梯度为0。

内容的提问来源于stack exchange,提问作者Hossam Tarek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 00:55:24