You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Keras/TensorFlow中构建优化ROC AUC的自定义损失函数

问题:自定义负ROC AUC损失函数时出现"No gradients provided for any variable"报错

我正在开发二分类预测的Jupyter Notebook,用Keras/TensorFlow实现概率预测,核心目标是优化ROC AUC指标。模型定义代码如下:

train_x, train_y = df[predictors].iloc[train_idx].values, df['target'].iloc[train_idx].values
valid_x, valid_y = df[predictors].iloc[valid_idx].values, df['target'].iloc[valid_idx].values
train_x = np.asarray(train_x).astype('float32')
train_y = np.asarray(train_y).astype('float32')
valid_x = np.asarray(valid_x).astype('float32')
valid_y = np.asarray(valid_y).astype('float32')
# Define FNN model architecture
num_features = train_x.shape[1]
clf = Sequential([
    Dense(128, activation='relu', input_shape=(num_features,)),
    Dense(64, activation='relu'),
    Dense(32, activation='relu'),
    Dense(1, activation='sigmoid')
])

# Compile the model with custom loss function and ROC AUC metric
clf.compile(optimizer='adam',
            loss=custom_loss,
            metrics=[tf.keras.metrics.AUC()])

# Define early stopping callback
early_stopping = EarlyStopping(monitor='val_loss', patience=5, restore_best_weights=True)
# Define learning rate scheduler
lr_scheduler = ReduceLROnPlateau(monitor='val_loss', factor=0.1, patience=3, verbose=1)

# Train the model
history = clf.fit(
    train_x, train_y, 
    epochs=50, 
    batch_size=32, 
    validation_data=(valid_x, valid_y),
    callbacks=[early_stopping, lr_scheduler],
    verbose=1
)

尝试定义以负ROC AUC为损失的函数:

def custom_loss(y_true, y_pred):
    y_true = tf.squeeze(y_true, axis=1)
    y_pred  = tf.squeeze(y_pred, axis=1)

    # Calculate ROC AUC score
    m = tf.keras.metrics.AUC()
    m.update_state(y_true, y_pred)

    # Take negative of ROC AUC as loss
    print("result: ",m.result())
    return -(m.result())

运行时触发报错:

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
Cell In[39], line 5
      3 pd.set_option('display.max_columns', 100)
      4 with timer("Pipeline total time"):
----> 5     main(debug= False)

Cell In[4], line 165, in main(debug)
    163     predictors = list(filter(lambda v: v not in del_features, df.columns))
    164     cat_cols = list(df.select_dtypes("object").columns)
--> 165     model = kfold_lightgbm_sklearn(df, cat_cols)
    171 #with timer("Feature importance assesment"):
    172     
    173  #   get_features_importances(predictors, model)
    178 with timer("Submission"):

Cell In[18], line 100, in kfold_lightgbm_sklearn(***failed resolving arguments***)
     97 lr_scheduler = ReduceLROnPlateau(monitor='val_loss', factor=0.1, patience=3, verbose=1)
     99 # Train the model
--> 100 history = clf.fit(train_x, train_y, 
    101                   epochs=50, 
    102                   batch_size=32, 
    103                   validation_data=(valid_x, valid_y),
    104                   callbacks=[early_stopping, lr_scheduler],
    105                   verbose=1)
    108 fitted_models.append(clf)
    110 if EVALUATE_VALIDATION_SET:
    111     # Obtain probabilities for each class

File /opt/conda/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:123, in filter_traceback.<locals>.error_handler(*args, **kwargs)
    120     filtered_tb = _process_traceback_frames(e.__traceback__)
    121     # To get the full stack trace, call:
    122     # `keras.config.disable_traceback_filtering()`
--> 123     raise e.with_traceback(filtered_tb) from None
    124 finally:
    125     del filtered_tb

File /opt/conda/lib/python3.10/site-packages/keras/src/optimizers/base_optimizer.py:571, in BaseOptimizer._filter_empty_gradients(self, grads, vars)
    567 filtered = [
    568     (g, v) for g, v in zip(grads, vars) if g is not None
    569 ]
    570 if not filtered:
--> 571     raise ValueError("No gradients provided for any variable.")
    572 if len(filtered) < len(grads):
    573     missing_grad_vars = [
    574         v for g, v in zip(grads, vars) if g is None
    575     ]

ValueError: No gradients provided for any variable.

错误原因

直接用tf.keras.metrics.AUC()计算损失的核心问题是:

  • Keras的metrics类是为模型评估设计的,内部包含不可微分的操作(比如排序、阈值统计、累计更新等),TensorFlow无法对这些操作计算梯度。
  • 每次在损失函数中创建新的AUC实例,其内部状态与模型参数的计算图完全断开,导致梯度无法反向传播到模型变量。

解决方案

方案1:常规优化思路(推荐)

ROC AUC是评估指标,本身不适合作为损失函数。正确的做法是用二元交叉熵损失(二分类任务的标准损失)训练模型,同时将ROC AUC作为监控指标,验证模型性能。

调整代码如下:

# 编译时使用标准二元交叉熵损失,保留ROC AUC作为监控指标
clf.compile(optimizer='adam',
            loss='binary_crossentropy',
            metrics=[tf.keras.metrics.AUC(name='roc_auc')])

这种方案的优势是:二元交叉熵可微分,梯度计算稳定,且在大多数二分类场景下,优化交叉熵损失能间接提升ROC AUC指标。

方案2:近似ROC AUC的可微分损失(排序损失)

如果一定要直接优化与ROC AUC相关的目标,可以使用成对排序损失(Pairwise Ranking Loss)。ROC AUC的本质是衡量"正样本的预测概率大于负样本"的比例,排序损失通过惩罚正样本预测概率小于负样本的情况,近似实现这一目标,且完全可微分。

自定义排序损失的代码实现:

def ranking_loss(y_true, y_pred):
    # 分离正负样本的预测值
    y_pred_pos = tf.boolean_mask(y_pred, tf.cast(y_true, tf.bool))
    y_pred_neg = tf.boolean_mask(y_pred, ~tf.cast(y_true, tf.bool))
    
    # 计算所有正负样本对的损失:max(0, 1 - (pos_pred - neg_pred))
    pos_pred_expanded = tf.expand_dims(y_pred_pos, axis=1)
    neg_pred_expanded = tf.expand_dims(y_pred_neg, axis=0)
    loss = tf.maximum(0.0, 1.0 - (pos_pred_expanded - neg_pred_expanded))
    
    # 对所有样本对的损失取平均
    return tf.reduce_mean(loss)

然后在编译时使用这个损失:

clf.compile(optimizer='adam',
            loss=ranking_loss,
            metrics=[tf.keras.metrics.AUC(name='roc_auc')])

注意:排序损失对样本分布敏感,如果正负样本数量差异极大,可能需要调整损失计算方式(比如加权)。


内容的提问来源于stack exchange,提问作者Matous Famera

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 17:22:05