自定义CSI损失函数导致GradientTape返回None的问题求助
自定义CSI损失函数导致梯度全为None的问题解决
问题概述
在TensorFlow中搭建处理64x64像素图像的简单CNN时,使用自定义CSI(Critical Success Index)损失函数,训练过程中梯度始终返回全为None的列表,导致后续优化步骤失败。使用keras.losses.BinaryCrossentropy等标准二分类损失函数时代码可正常运行,且已确认y_true和y_pred的尺寸、数据类型完全一致。
相关代码
自定义损失函数
from keras import backend as K @tf.function def custom_csi_loss(y_true, y_pred): # Define the target class target_class = 1 # Calculate the true positives, false positives, and false negatives true_positives = K.sum(K.round(K.clip(y_true * y_pred, 0, 1))) false_positives = K.sum(K.round(K.clip(y_pred - y_true, 0, 1))) false_negatives = K.sum(K.round(K.clip(y_true - y_pred, 0, 1))) # Calculate the CSI csi = true_positives / (true_positives + false_negatives + false_positives) # Return the negative of the CSI as the loss (since we want to minimize the loss) return -csi
模型定义
def build_scnn(shape=(128, 128, 3), k_init="he_normal", dilation_rate=(1, 1), dtype=tf.float32): inputs = Input(shape=shape) normalized = BatchNormalization(axis=3)(inputs) x = Conv2D(64, 3, padding="same", activation="relu", kernel_initializer=k_init)(normalized) x = Conv2D(128, 3, padding="same", activation="relu", dilation_rate=dilation_rate, kernel_initializer=k_init)(x) x = Conv2D(128, 3, padding="same", activation="relu", kernel_initializer=k_init)(x) outputs = Conv2D(1, 1, padding="same", activation="sigmoid", dtype=dtype)(x) outputs = Reshape((64 * 64, 1))(outputs) scnn = Model(inputs, outputs, name="SCNN") return scnn scnn = build_scnn(shape=(64, 64, len(gdf[features].columns)), k_init=k_init, dilation_rate=dilation_rate)
训练步骤
@tf.function def train_step(x, y): with tf.GradientTape(watch_accessed_variables=True) as tape: tape.watch(scnn.trainable_variables) y_pred = scnn(x, training=True) loss = loss_fn(y, y_pred) gradients = tape.gradient(loss, scnn.trainable_variables) # differentiate loss wrt scnn weights print(f"gradients: {gradients}") optimizer.apply_gradients(zip(gradients, scnn.trainable_variables)) return loss, y_pred
训练主循环
for epoch in range(epochs): epoch_loss = 0 epoch_csi = 0 num_batches = 0 for x, y, w in train.map(weight_func): y = tf.cast(y, dtype=tf.float32) loss, y_pred = train_step(x, y) epoch_loss += loss epoch_csi += metrics[0](y, y_pred) num_batches += 1 epoch_loss /= num_batches epoch_csi /= num_batches
问题原因
自定义损失函数中的K.round操作是不可微分的阶跃函数:在非整数点导数为0,整数点导数不存在。TensorFlow无法通过该操作计算梯度,导致梯度无法传递回模型的可训练参数,最终所有梯度返回None。
解决方案
替换不可微分的round操作,使用连续可微分的近似计算,同时添加小epsilon避免除以0的情况:
方案1:用陡峭sigmoid近似round操作
import tensorflow as tf from keras import backend as K @tf.function def custom_csi_loss(y_true, y_pred): # 用陡峭的sigmoid函数近似round操作,保证可微分 def steep_sigmoid(x): return tf.sigmoid(x * 100) # 100控制陡峭程度,可根据需求调整 # 计算连续版本的TP、FP、FN true_positives = K.sum(steep_sigmoid(y_true * y_pred) * K.clip(y_true * y_pred, 0, 1)) false_positives = K.sum(K.clip(y_pred - y_true, 0, 1)) false_negatives = K.sum(K.clip(y_true - y_pred, 0, 1)) # 添加epsilon避免分母为0 epsilon = K.epsilon() denominator = true_positives + false_negatives + false_positives + epsilon csi = true_positives / denominator return -csi
方案2:直接使用连续值计算(简化版)
from keras import backend as K @tf.function def custom_csi_loss(y_true, y_pred): true_positives = K.sum(K.clip(y_true * y_pred, 0, 1)) false_positives = K.sum(K.clip(y_pred - y_true, 0, 1)) false_negatives = K.sum(K.clip(y_true - y_pred, 0, 1)) epsilon = K.epsilon() csi = true_positives / (true_positives + false_negatives + false_positives + epsilon) return -csi
说明
- 陡峭sigmoid的作用是在保留近似离散判断的同时,保证函数的可微性,让梯度能够正常传递。
- 添加
K.epsilon()是为了防止分母为0导致NaN,保证损失计算的稳定性。
内容的提问来源于stack exchange,提问作者ThreeOrangeOneRed
相关产品推荐
相关产品推荐

