You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义CSI损失函数导致GradientTape返回None的问题求助

自定义CSI损失函数导致梯度全为None的问题解决

问题概述

在TensorFlow中搭建处理64x64像素图像的简单CNN时,使用自定义CSI(Critical Success Index)损失函数,训练过程中梯度始终返回全为None的列表,导致后续优化步骤失败。使用keras.losses.BinaryCrossentropy等标准二分类损失函数时代码可正常运行,且已确认y_true和y_pred的尺寸、数据类型完全一致。

相关代码

自定义损失函数

from keras import backend as K

@tf.function
def custom_csi_loss(y_true, y_pred):
    # Define the target class
    target_class = 1

    # Calculate the true positives, false positives, and false negatives
    true_positives = K.sum(K.round(K.clip(y_true * y_pred, 0, 1)))
    false_positives = K.sum(K.round(K.clip(y_pred - y_true, 0, 1)))
    false_negatives = K.sum(K.round(K.clip(y_true - y_pred, 0, 1)))

    # Calculate the CSI
    csi = true_positives / (true_positives + false_negatives + false_positives)

    # Return the negative of the CSI as the loss (since we want to minimize the loss)
    return -csi

模型定义

def build_scnn(shape=(128, 128, 3), k_init="he_normal", dilation_rate=(1, 1), dtype=tf.float32):
    inputs = Input(shape=shape)
    normalized = BatchNormalization(axis=3)(inputs)

    x = Conv2D(64, 3, padding="same", activation="relu", kernel_initializer=k_init)(normalized)
    x = Conv2D(128, 3, padding="same", activation="relu", dilation_rate=dilation_rate, kernel_initializer=k_init)(x)
    x = Conv2D(128, 3, padding="same", activation="relu", kernel_initializer=k_init)(x)

    outputs = Conv2D(1, 1, padding="same", activation="sigmoid", dtype=dtype)(x)
    outputs = Reshape((64 * 64, 1))(outputs)
    scnn = Model(inputs, outputs, name="SCNN")
    return scnn

scnn = build_scnn(shape=(64, 64, len(gdf[features].columns)),
                  k_init=k_init,
                  dilation_rate=dilation_rate)

训练步骤

@tf.function
def train_step(x, y):
    with tf.GradientTape(watch_accessed_variables=True) as tape:
        tape.watch(scnn.trainable_variables)
        y_pred = scnn(x, training=True)
        loss = loss_fn(y, y_pred)
    gradients = tape.gradient(loss, scnn.trainable_variables)  # differentiate loss wrt scnn weights
    print(f"gradients: {gradients}")
    optimizer.apply_gradients(zip(gradients, scnn.trainable_variables))
    return loss, y_pred

训练主循环

for epoch in range(epochs):
    epoch_loss = 0
    epoch_csi = 0
    num_batches = 0
    for x, y, w in train.map(weight_func):
        y = tf.cast(y, dtype=tf.float32)
        loss, y_pred = train_step(x, y)
        epoch_loss += loss
        epoch_csi += metrics[0](y, y_pred)
        num_batches += 1
    
    epoch_loss /= num_batches
    epoch_csi /= num_batches

问题原因

自定义损失函数中的K.round操作是不可微分的阶跃函数:在非整数点导数为0,整数点导数不存在。TensorFlow无法通过该操作计算梯度,导致梯度无法传递回模型的可训练参数,最终所有梯度返回None。

解决方案

替换不可微分的round操作,使用连续可微分的近似计算,同时添加小epsilon避免除以0的情况:

方案1:用陡峭sigmoid近似round操作

import tensorflow as tf
from keras import backend as K

@tf.function
def custom_csi_loss(y_true, y_pred):
    # 用陡峭的sigmoid函数近似round操作,保证可微分
    def steep_sigmoid(x):
        return tf.sigmoid(x * 100)  # 100控制陡峭程度,可根据需求调整
    
    # 计算连续版本的TP、FP、FN
    true_positives = K.sum(steep_sigmoid(y_true * y_pred) * K.clip(y_true * y_pred, 0, 1))
    false_positives = K.sum(K.clip(y_pred - y_true, 0, 1))
    false_negatives = K.sum(K.clip(y_true - y_pred, 0, 1))
    
    # 添加epsilon避免分母为0
    epsilon = K.epsilon()
    denominator = true_positives + false_negatives + false_positives + epsilon
    csi = true_positives / denominator
    
    return -csi

方案2:直接使用连续值计算(简化版)

from keras import backend as K

@tf.function
def custom_csi_loss(y_true, y_pred):
    true_positives = K.sum(K.clip(y_true * y_pred, 0, 1))
    false_positives = K.sum(K.clip(y_pred - y_true, 0, 1))
    false_negatives = K.sum(K.clip(y_true - y_pred, 0, 1))
    
    epsilon = K.epsilon()
    csi = true_positives / (true_positives + false_negatives + false_positives + epsilon)
    
    return -csi

说明

  • 陡峭sigmoid的作用是在保留近似离散判断的同时,保证函数的可微性,让梯度能够正常传递。
  • 添加K.epsilon()是为了防止分母为0导致NaN,保证损失计算的稳定性。

内容的提问来源于stack exchange,提问作者ThreeOrangeOneRed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 07:35:20