You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用tf.GradientTape()让神经网络第二层变量不可训练?

解决TensorFlow中第二层网络参数不可训练的问题

针对你的需求,有三种简单可行的方法实现第二层参数(w2、b2)不可训练,同时保证第一层参数正常更新:

方法一:定义变量时设置trainable=False

TensorFlow中tf.Variable的trainable参数默认是True,直接将w2、b2的该参数设为False,GradientTape会自动忽略这些变量的梯度计算。注意需要将原本的张量转换为tf.Variable(否则无法被优化器更新):

# 第一层参数:可训练
w1 = tf.Variable(tf.random.truncated_normal([28*28, 256]), trainable=True)
b1 = tf.Variable(tf.zeros([256]), trainable=True)
# 第二层参数:不可训练
w2 = tf.Variable(tf.random.truncated_normal([256, 10]), trainable=False)  # 修正原代码形状不匹配问题:w2应与b2维度对应
b2 = tf.Variable(tf.zeros([10]), trainable=False)

后续训练时,梯度只会计算w1、b1的部分,优化器也只会更新这两个变量。

方法二:用tf.stop_gradient阻断梯度传播

如果不想修改变量定义,可以在计算第二层输出时,用tf.stop_gradient包裹w2和b2,阻止梯度反向传播到这些参数:

for (x,y) in db:
    x = tf.reshape(x, [-1, 28*28])
    with tf.GradientTape() as tape:
        h1 = x@w1 + tf.broadcast_to(b1, [x.shape[0], 256])
        h1 = tf.nn.relu(h1)
        # 阻断w2、b2的梯度传播
        h2 = h1@tf.stop_gradient(w2) + tf.broadcast_to(tf.stop_gradient(b2), [x.shape[0], 10])
        out = tf.nn.relu(h2)
        y_onehot = tf.one_hot(y, depth=10)
        loss = tf.square(y_onehot - out)
        loss = tf.reduce_mean(loss)
    # 计算梯度时仅得到w1、b1的梯度
    grads = tape.gradient(loss, [w1, b1])
    # 执行参数更新(示例用SGD优化器)
    optimizer.apply_gradients(zip(grads, [w1, b1]))

方法三:手动指定求导变量

在调用tape.gradient时,只传入需要计算梯度的变量(w1、b1),即使GradientTape跟踪了w2、b2,也不会计算它们的梯度:

for (x,y) in db:
    x = tf.reshape(x, [-1, 28*28])
    with tf.GradientTape() as tape:
        h1 = x@w1 + tf.broadcast_to(b1, [x.shape[0], 256])
        h1 = tf.nn.relu(h1)
        h2 = h1@w2 + tf.broadcast_to(b2, [x.shape[0], 10])
        out = tf.nn.relu(h2)
        y_onehot = tf.one_hot(y, depth=10)
        loss = tf.square(y_onehot - out)
        loss = tf.reduce_mean(loss)
    # 仅对w1、b1计算梯度
    grads = tape.gradient(loss, [w1, b1])
    # 只更新w1、b1
    optimizer.apply_gradients(zip(grads, [w1, b1]))

注意事项

原代码中存在形状不匹配的问题:w2的形状是[256,50],但b2是[10],这会导致h1@w2的输出形状([batch_size,50])与b2无法广播相加,建议将w2的形状修正为[256,10]。

内容的提问来源于stack exchange,提问作者David H. J.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 22:32:20