如何在Keras/TensorFlow中构建优化ROC AUC的自定义损失函数
问题:自定义负ROC AUC损失函数时出现"No gradients provided for any variable"报错
我正在开发二分类预测的Jupyter Notebook,用Keras/TensorFlow实现概率预测,核心目标是优化ROC AUC指标。模型定义代码如下:
train_x, train_y = df[predictors].iloc[train_idx].values, df['target'].iloc[train_idx].values valid_x, valid_y = df[predictors].iloc[valid_idx].values, df['target'].iloc[valid_idx].values train_x = np.asarray(train_x).astype('float32') train_y = np.asarray(train_y).astype('float32') valid_x = np.asarray(valid_x).astype('float32') valid_y = np.asarray(valid_y).astype('float32') # Define FNN model architecture num_features = train_x.shape[1] clf = Sequential([ Dense(128, activation='relu', input_shape=(num_features,)), Dense(64, activation='relu'), Dense(32, activation='relu'), Dense(1, activation='sigmoid') ]) # Compile the model with custom loss function and ROC AUC metric clf.compile(optimizer='adam', loss=custom_loss, metrics=[tf.keras.metrics.AUC()]) # Define early stopping callback early_stopping = EarlyStopping(monitor='val_loss', patience=5, restore_best_weights=True) # Define learning rate scheduler lr_scheduler = ReduceLROnPlateau(monitor='val_loss', factor=0.1, patience=3, verbose=1) # Train the model history = clf.fit( train_x, train_y, epochs=50, batch_size=32, validation_data=(valid_x, valid_y), callbacks=[early_stopping, lr_scheduler], verbose=1 )
尝试定义以负ROC AUC为损失的函数:
def custom_loss(y_true, y_pred): y_true = tf.squeeze(y_true, axis=1) y_pred = tf.squeeze(y_pred, axis=1) # Calculate ROC AUC score m = tf.keras.metrics.AUC() m.update_state(y_true, y_pred) # Take negative of ROC AUC as loss print("result: ",m.result()) return -(m.result())
运行时触发报错:
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) Cell In[39], line 5 3 pd.set_option('display.max_columns', 100) 4 with timer("Pipeline total time"): ----> 5 main(debug= False) Cell In[4], line 165, in main(debug) 163 predictors = list(filter(lambda v: v not in del_features, df.columns)) 164 cat_cols = list(df.select_dtypes("object").columns) --> 165 model = kfold_lightgbm_sklearn(df, cat_cols) 171 #with timer("Feature importance assesment"): 172 173 # get_features_importances(predictors, model) 178 with timer("Submission"): Cell In[18], line 100, in kfold_lightgbm_sklearn(***failed resolving arguments***) 97 lr_scheduler = ReduceLROnPlateau(monitor='val_loss', factor=0.1, patience=3, verbose=1) 99 # Train the model --> 100 history = clf.fit(train_x, train_y, 101 epochs=50, 102 batch_size=32, 103 validation_data=(valid_x, valid_y), 104 callbacks=[early_stopping, lr_scheduler], 105 verbose=1) 108 fitted_models.append(clf) 110 if EVALUATE_VALIDATION_SET: 111 # Obtain probabilities for each class File /opt/conda/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py:123, in filter_traceback.<locals>.error_handler(*args, **kwargs) 120 filtered_tb = _process_traceback_frames(e.__traceback__) 121 # To get the full stack trace, call: 122 # `keras.config.disable_traceback_filtering()` --> 123 raise e.with_traceback(filtered_tb) from None 124 finally: 125 del filtered_tb File /opt/conda/lib/python3.10/site-packages/keras/src/optimizers/base_optimizer.py:571, in BaseOptimizer._filter_empty_gradients(self, grads, vars) 567 filtered = [ 568 (g, v) for g, v in zip(grads, vars) if g is not None 569 ] 570 if not filtered: --> 571 raise ValueError("No gradients provided for any variable.") 572 if len(filtered) < len(grads): 573 missing_grad_vars = [ 574 v for g, v in zip(grads, vars) if g is None 575 ] ValueError: No gradients provided for any variable.
错误原因
直接用tf.keras.metrics.AUC()计算损失的核心问题是:
- Keras的metrics类是为模型评估设计的,内部包含不可微分的操作(比如排序、阈值统计、累计更新等),TensorFlow无法对这些操作计算梯度。
- 每次在损失函数中创建新的
AUC实例,其内部状态与模型参数的计算图完全断开,导致梯度无法反向传播到模型变量。
解决方案
方案1:常规优化思路(推荐)
ROC AUC是评估指标,本身不适合作为损失函数。正确的做法是用二元交叉熵损失(二分类任务的标准损失)训练模型,同时将ROC AUC作为监控指标,验证模型性能。
调整代码如下:
# 编译时使用标准二元交叉熵损失,保留ROC AUC作为监控指标 clf.compile(optimizer='adam', loss='binary_crossentropy', metrics=[tf.keras.metrics.AUC(name='roc_auc')])
这种方案的优势是:二元交叉熵可微分,梯度计算稳定,且在大多数二分类场景下,优化交叉熵损失能间接提升ROC AUC指标。
方案2:近似ROC AUC的可微分损失(排序损失)
如果一定要直接优化与ROC AUC相关的目标,可以使用成对排序损失(Pairwise Ranking Loss)。ROC AUC的本质是衡量"正样本的预测概率大于负样本"的比例,排序损失通过惩罚正样本预测概率小于负样本的情况,近似实现这一目标,且完全可微分。
自定义排序损失的代码实现:
def ranking_loss(y_true, y_pred): # 分离正负样本的预测值 y_pred_pos = tf.boolean_mask(y_pred, tf.cast(y_true, tf.bool)) y_pred_neg = tf.boolean_mask(y_pred, ~tf.cast(y_true, tf.bool)) # 计算所有正负样本对的损失:max(0, 1 - (pos_pred - neg_pred)) pos_pred_expanded = tf.expand_dims(y_pred_pos, axis=1) neg_pred_expanded = tf.expand_dims(y_pred_neg, axis=0) loss = tf.maximum(0.0, 1.0 - (pos_pred_expanded - neg_pred_expanded)) # 对所有样本对的损失取平均 return tf.reduce_mean(loss)
然后在编译时使用这个损失:
clf.compile(optimizer='adam', loss=ranking_loss, metrics=[tf.keras.metrics.AUC(name='roc_auc')])
注意:排序损失对样本分布敏感,如果正负样本数量差异极大,可能需要调整损失计算方式(比如加权)。
内容的提问来源于stack exchange,提问作者Matous Famera
相关产品推荐
相关产品推荐

