使用GradientTape时遭遇“No gradients provided for any variable”错误求助
问题描述
首次尝试用TensorFlow的GradientTape实现自定义复杂损失函数训练CNN,出现ValueError: No gradients provided for any variable错误。使用标准损失函数时模型可正常收敛,相关代码及报错信息如下:
可复现代码
# imports import numpy as np import tensorflow as tf import sklearn from tensorflow import keras from tensorflow.keras import layers from sklearn.model_selection import train_test_split from sklearn.svm import SVC tf.config.run_functions_eagerly(True) #my loss function def my_loss_fn(y_true, y_pred): # train SVM classifier VarC=1E6 VarGamma='scale' clf = SVC(kernel='rbf', C=VarC, gamma=VarGamma, probability=True ) clf.fit(y_pred, y_true) y_pred = clf.predict_proba(y_pred) scce = tf.keras.losses.SparseCategoricalCrossentropy() return scce(y_true, y_pred) #creating inputs to demontration X0=0.5*np.ones((12,12)) X0[2:12:4,:]=0 X0[3:12:4,:]=0 X1=0.5*np.ones((12,12)) X1[1:12:4,:]=0 X1[2:12:4,:]=0 X1=np.transpose(X1) X=np.zeros((2000,12,12)) for i in range(0,1000): X[i]=X0+np.random.rand(12,12) for i in range(1000,2000): X[i]=X1+np.random.rand(12,12) y=np.zeros(2000, dtype=int) y[1000:2000]=1 x_train, x_val, y_train, y_val = train_test_split(X, y, train_size=0.5) x_val, x_test, y_val, y_test = train_test_split(x_val, y_val, train_size=0.5) x_train = tf.convert_to_tensor(x_train) x_val = tf.convert_to_tensor(x_val) x_test = tf.convert_to_tensor(x_test) y_train = tf.convert_to_tensor(y_train) y_val = tf.convert_to_tensor(y_val) y_test = tf.convert_to_tensor(y_test) inputs = keras.Input((12,12,1), name='images') x0 = tf.keras.layers.Conv2D(8,4,strides=4)(inputs) x0 = tf.keras.layers.AveragePooling2D(pool_size=(3, 3), name='pooling')(x0) outputs = tf.keras.layers.Flatten(name='predictions')(x0) model = keras.Model(inputs=inputs, outputs=outputs) optimizer=tf.keras.optimizers.Adam(learning_rate=0.001) # Instantiate a loss function. loss_fn = my_loss_fn # Prepare the training dataset. batch_size = 256 train_dataset = tf.data.Dataset.from_tensor_slices((x_train, y_train)) train_dataset = train_dataset.shuffle(buffer_size=1024).batch(batch_size) epochs = 100 for epoch in range(epochs): print('Start of epoch %d' % (epoch,)) # Iterate over the batches of the dataset. for step, (x_batch_train, y_batch_train) in enumerate(train_dataset): # Open a GradientTape to record the operations run # during the forward pass, which enables autodifferentiation. with tf.GradientTape() as tape: tape.watch(model.trainable_weights) # Run the forward pass of the layer. # The operations that the layer applies # to its inputs are going to be recorded # on the GradientTape. logits = model(x_batch_train, training=True) # Logits for this minibatch # Compute the loss value for this minibatch. loss_value = loss_fn(y_batch_train, logits) # Use the gradient tape to automatically retrieve # the gradients of the trainable variables with respect to the loss. grads = tape.gradient(loss_value, model.trainable_weights) # Run one step of gradient descent by updating # the value of the variables to minimize the loss. optimizer.apply_gradients(zip(grads, model.trainable_weights)) # Log every 200 batches. if step % 200 == 0: print('Training loss (for one batch) at step %s: %s' % (step, float(loss_value))) print('Seen so far: %s samples' % ((step + 1) * 64))
报错信息
ValueError: No gradients provided for any variable: (['conv2d_2/kernel:0', 'conv2d_2/bias:0'],). Provided
grads_and_varsis ((None, <tf.Variable 'conv2d_2/kernel:0' shape=(4, 4, 1, 8) dtype=float32, nump
标准损失函数下的正常代码片段
inputs = keras.Input((12,12,1), name='images') x0 = tf.keras.layers.Conv2D(8,4,strides=4)(inputs) x0 = tf.keras.layers.AveragePooling2D(pool_size=(3, 3), name='pooling')(x0) x0 = tf.keras.layers.Flatten(name='features')(x0) x0 = layers.Dense(16, name='meta_features')(x0) outputs = layers.Dense(2, name='predictions')(x0) model = keras.Model(inputs=inputs, outputs=outputs) loss_fn = keras.losses.SparseCategoricalCrossentropy(from_logits=True)
问题根源
错误的核心原因是自定义损失函数中使用了Scikit-learn的SVC模型:
- TensorFlow的GradientTape仅能追踪TensorFlow原生操作或兼容AutoGraph的代码的梯度,而Scikit-learn的SVC是基于纯Python/Numpy实现的非可微分操作。
- 在
my_loss_fn中,clf.fit()和clf.predict_proba()完全脱离了TensorFlow的计算图,导致损失值与CNN模型可训练参数之间的梯度链路断裂,最终tape.gradient()返回None,触发报错。
解决方法
方法1:用TensorFlow原生操作重写SVM逻辑
如果必须保留损失中的SVM逻辑,需将SVM的训练和预测用TensorFlow可微分操作实现。以RBF核SVM为例,可近似实现如下:
def my_loss_fn(y_true, y_pred): # 将y_true转为{-1,1}格式适配SVM y_true = tf.cast(y_true, tf.float32) * 2 - 1 # 计算RBF核:exp(-gamma*||x_i - x_j||²) gamma = tf.convert_to_tensor(1.0 / (y_pred.shape[-1] * tf.math.reduce_variance(y_pred))) pairwise_dist = tf.norm(tf.expand_dims(y_pred, 1) - tf.expand_dims(y_pred, 0), axis=2) ** 2 kernel = tf.exp(-gamma * pairwise_dist) # SVM对偶问题损失近似(简化版) alpha = tf.Variable(tf.zeros_like(y_true), trainable=True) loss_svm = tf.reduce_sum(alpha) - 0.5 * tf.reduce_sum(tf.matmul(tf.matmul(tf.diag(y_true), kernel), tf.diag(y_true)) * tf.matmul(alpha[:, None], alpha[None, :])) loss_svm += 1e6 * tf.reduce_sum(tf.square(alpha)) # 对应C参数的正则项 # 结合交叉熵损失 y_pred_proba = tf.nn.softmax(y_pred) scce = tf.keras.losses.SparseCategoricalCrossentropy() loss_scce = scce(tf.cast((y_true + 1)/2, tf.int32), y_pred_proba) return loss_svm + loss_scce
注:这是简化的近似实现,实际使用需根据需求调整SVM损失计算逻辑,确保全链路可微分。
方法2:拆分训练流程,避免损失内训练其他模型
将SVM作为后处理模块,不嵌入损失函数:
- 先训练CNN提取特征,固定CNN参数;
- 用CNN提取的特征训练SVM分类器;
- 若需端到端训练,可将SVM参数转为TensorFlow可训练变量,用可微分方式优化。
方法3:替换为可微分的替代方案
放弃在损失中使用SVM,改用TensorFlow原生可微分分类模块(如Dense层+Softmax),即采用你提供的标准损失函数代码逻辑,或用其他可微分度量替代SVM逻辑。
内容的提问来源于stack exchange,提问作者Rehal
相关产品推荐
相关产品推荐

