You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用GradientTape时遭遇“No gradients provided for any variable”错误求助

问题描述

首次尝试用TensorFlow的GradientTape实现自定义复杂损失函数训练CNN,出现ValueError: No gradients provided for any variable错误。使用标准损失函数时模型可正常收敛,相关代码及报错信息如下:

可复现代码

# imports
import numpy as np
import tensorflow as tf
import sklearn

from tensorflow import keras
from tensorflow.keras import layers
from sklearn.model_selection import train_test_split
from sklearn.svm import SVC
tf.config.run_functions_eagerly(True)
#my loss function
def my_loss_fn(y_true, y_pred):
    # train SVM classifier
    VarC=1E6
    VarGamma='scale'
    clf = SVC(kernel='rbf', C=VarC, gamma=VarGamma, probability=True )
    clf.fit(y_pred, y_true)
    y_pred = clf.predict_proba(y_pred)
    scce = tf.keras.losses.SparseCategoricalCrossentropy()
    return scce(y_true, y_pred)
#creating inputs to demontration
X0=0.5*np.ones((12,12))
X0[2:12:4,:]=0
X0[3:12:4,:]=0
X1=0.5*np.ones((12,12))
X1[1:12:4,:]=0
X1[2:12:4,:]=0
X1=np.transpose(X1)
X=np.zeros((2000,12,12))

for i in range(0,1000):
    X[i]=X0+np.random.rand(12,12)
for i in range(1000,2000):
    X[i]=X1+np.random.rand(12,12)
y=np.zeros(2000, dtype=int)
y[1000:2000]=1
x_train, x_val, y_train, y_val = train_test_split(X, y, train_size=0.5)
x_val, x_test, y_val, y_test = train_test_split(x_val, y_val, train_size=0.5)
x_train = tf.convert_to_tensor(x_train)
x_val = tf.convert_to_tensor(x_val)
x_test = tf.convert_to_tensor(x_test)
y_train = tf.convert_to_tensor(y_train)
y_val = tf.convert_to_tensor(y_val)
y_test = tf.convert_to_tensor(y_test)

inputs = keras.Input((12,12,1), name='images')
x0 = tf.keras.layers.Conv2D(8,4,strides=4)(inputs)
x0 = tf.keras.layers.AveragePooling2D(pool_size=(3, 3), name='pooling')(x0)
outputs = tf.keras.layers.Flatten(name='predictions')(x0)

model = keras.Model(inputs=inputs, outputs=outputs)
optimizer=tf.keras.optimizers.Adam(learning_rate=0.001)
# Instantiate a loss function.
loss_fn = my_loss_fn

# Prepare the training dataset.
batch_size = 256
train_dataset = tf.data.Dataset.from_tensor_slices((x_train, y_train))
train_dataset = train_dataset.shuffle(buffer_size=1024).batch(batch_size)

epochs = 100
for epoch in range(epochs):
    print('Start of epoch %d' % (epoch,))
    # Iterate over the batches of the dataset.
    for step, (x_batch_train, y_batch_train) in enumerate(train_dataset):
        # Open a GradientTape to record the operations run
        # during the forward pass, which enables autodifferentiation.
        with tf.GradientTape() as tape:
            tape.watch(model.trainable_weights)
            # Run the forward pass of the layer.
            # The operations that the layer applies
            # to its inputs are going to be recorded
            # on the GradientTape.
            logits = model(x_batch_train, training=True)  # Logits for this minibatch
            # Compute the loss value for this minibatch.
            loss_value = loss_fn(y_batch_train, logits)
        # Use the gradient tape to automatically retrieve
        # the gradients of the trainable variables with respect to the loss.
        grads = tape.gradient(loss_value, model.trainable_weights)
        # Run one step of gradient descent by updating
        # the value of the variables to minimize the loss.
        optimizer.apply_gradients(zip(grads, model.trainable_weights))
        # Log every 200 batches.
        if step % 200 == 0:
            print('Training loss (for one batch) at step %s: %s' % (step, float(loss_value)))
            print('Seen so far: %s samples' % ((step + 1) * 64))

报错信息

ValueError: No gradients provided for any variable: (['conv2d_2/kernel:0', 'conv2d_2/bias:0'],). Provided grads_and_vars is ((None, <tf.Variable 'conv2d_2/kernel:0' shape=(4, 4, 1, 8) dtype=float32, nump

标准损失函数下的正常代码片段

inputs = keras.Input((12,12,1), name='images')
x0 = tf.keras.layers.Conv2D(8,4,strides=4)(inputs)
x0 = tf.keras.layers.AveragePooling2D(pool_size=(3, 3), name='pooling')(x0)
x0 = tf.keras.layers.Flatten(name='features')(x0)
x0 = layers.Dense(16, name='meta_features')(x0)
outputs = layers.Dense(2, name='predictions')(x0)
model = keras.Model(inputs=inputs, outputs=outputs)
loss_fn = keras.losses.SparseCategoricalCrossentropy(from_logits=True)

问题根源

错误的核心原因是自定义损失函数中使用了Scikit-learn的SVC模型:

  • TensorFlow的GradientTape仅能追踪TensorFlow原生操作或兼容AutoGraph的代码的梯度,而Scikit-learn的SVC是基于纯Python/Numpy实现的非可微分操作。
  • 在my_loss_fn中,clf.fit()和clf.predict_proba()完全脱离了TensorFlow的计算图,导致损失值与CNN模型可训练参数之间的梯度链路断裂,最终tape.gradient()返回None,触发报错。

解决方法

方法1:用TensorFlow原生操作重写SVM逻辑

如果必须保留损失中的SVM逻辑,需将SVM的训练和预测用TensorFlow可微分操作实现。以RBF核SVM为例,可近似实现如下:

def my_loss_fn(y_true, y_pred):
    # 将y_true转为{-1,1}格式适配SVM
    y_true = tf.cast(y_true, tf.float32) * 2 - 1
    # 计算RBF核:exp(-gamma*||x_i - x_j||²)
    gamma = tf.convert_to_tensor(1.0 / (y_pred.shape[-1] * tf.math.reduce_variance(y_pred)))
    pairwise_dist = tf.norm(tf.expand_dims(y_pred, 1) - tf.expand_dims(y_pred, 0), axis=2) ** 2
    kernel = tf.exp(-gamma * pairwise_dist)
    
    # SVM对偶问题损失近似(简化版)
    alpha = tf.Variable(tf.zeros_like(y_true), trainable=True)
    loss_svm = tf.reduce_sum(alpha) - 0.5 * tf.reduce_sum(tf.matmul(tf.matmul(tf.diag(y_true), kernel), tf.diag(y_true)) * tf.matmul(alpha[:, None], alpha[None, :]))
    loss_svm += 1e6 * tf.reduce_sum(tf.square(alpha))  # 对应C参数的正则项
    
    # 结合交叉熵损失
    y_pred_proba = tf.nn.softmax(y_pred)
    scce = tf.keras.losses.SparseCategoricalCrossentropy()
    loss_scce = scce(tf.cast((y_true + 1)/2, tf.int32), y_pred_proba)
    
    return loss_svm + loss_scce

注:这是简化的近似实现,实际使用需根据需求调整SVM损失计算逻辑,确保全链路可微分。

方法2:拆分训练流程,避免损失内训练其他模型

将SVM作为后处理模块,不嵌入损失函数:

  1. 先训练CNN提取特征,固定CNN参数;
  2. 用CNN提取的特征训练SVM分类器;
  3. 若需端到端训练,可将SVM参数转为TensorFlow可训练变量,用可微分方式优化。

方法3:替换为可微分的替代方案

放弃在损失中使用SVM,改用TensorFlow原生可微分分类模块(如Dense层+Softmax),即采用你提供的标准损失函数代码逻辑,或用其他可微分度量替代SVM逻辑。


内容的提问来源于stack exchange,提问作者Rehal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 10:25:38