基于Keras的多标签图像分类自定义指标技术咨询
Got it, let's break down how to build custom Keras metrics for your 39-label multi-image classification task. Since each image can have multiple features, standard single-class metrics won’t cut it—we need metrics tailored to multi-label scenarios. Below are practical, ready-to-use custom metrics that align with your data setup (images shaped (1814,204,204,3) and labels shaped (1814,39)).
First, a quick reminder: for multi-label tasks, always use binary_crossentropy as your loss function (each label is an independent binary classification problem). Now let's dive into the metrics:
1. Hamming Loss (Lower = Better)
This metric measures the fraction of misclassified labels across all samples and labels—it’s perfect for getting a holistic view of how often your model gets any label wrong.
import tensorflow as tf from tensorflow.keras import backend as K def hamming_loss(y_true, y_pred): # Threshold predictions to binary 0/1 using 0.5 as the cutoff y_pred_bin = K.round(K.clip(y_pred, 0, 1)) # Count mismatched labels (false positives + false negatives) mismatches = K.sum(K.abs(y_true - y_pred_bin)) # Total number of label positions across all samples total_labels = K.cast(K.prod(K.shape(y_true)), dtype=K.floatx()) return mismatches / total_labels
2. Multi-Label Precision (Higher = Better)
Precision tells you how many of the labels your model predicted as positive are actually positive—great if you want to minimize false alarms (e.g., not flagging a feature that isn’t present).
def multi_label_precision(y_true, y_pred): y_pred_bin = K.round(K.clip(y_pred, 0, 1)) # True positives: labels predicted positive that are actually positive true_positives = K.sum(K.round(y_true * y_pred_bin)) # Total predicted positives: all labels the model flagged as positive predicted_positives = K.sum(y_pred_bin) # Avoid division by zero if no positives are predicted in a batch return K.switch(K.equal(predicted_positives, 0), 0.0, true_positives / predicted_positives)
3. Multi-Label Recall (Higher = Better)
Recall measures how many of the actual positive labels your model correctly identified—critical if you don’t want to miss important features (e.g., failing to detect a key feature that’s present).
def multi_label_recall(y_true, y_pred): y_pred_bin = K.round(K.clip(y_pred, 0, 1)) true_positives = K.sum(K.round(y_true * y_pred_bin)) # Total actual positives: all labels that are truly present in the data actual_positives = K.sum(y_true) # Avoid division by zero if no true positives exist in a batch return K.switch(K.equal(actual_positives, 0), 0.0, true_positives / actual_positives)
4. Multi-Label F1-Score (Higher = Better)
F1 is the harmonic mean of precision and recall—it balances both metrics, making it ideal when you care about avoiding both false positives and missed true positives.
def multi_label_f1(y_true, y_pred): precision = multi_label_precision(y_true, y_pred) recall = multi_label_recall(y_true, y_pred) # Calculate F1: 2*(P*R)/(P+R) return K.switch(K.equal(precision + recall, 0), 0.0, 2 * ((precision * recall) / (precision + recall)))
5. Mean Average Precision (mAP) (Higher = Better)
mAP is a more advanced metric that computes precision for each label at different recall thresholds, then averages those scores. It gives you a detailed view of performance across all 39 labels.
import numpy as np def mean_average_precision(y_true, y_pred): # Convert tensors to numpy arrays (works with TensorFlow 2.x eager execution) y_true_np = K.eval(y_true) y_pred_np = K.eval(y_pred) ap_scores = [] for label_idx in range(y_true_np.shape[1]): # Extract true and predicted values for the current label true = y_true_np[:, label_idx] pred = y_pred_np[:, label_idx] # Sort predictions in descending order of confidence sorted_indices = np.argsort(pred)[::-1] sorted_true = true[sorted_indices] # Calculate cumulative true positives and false positives tp = np.cumsum(sorted_true) fp = np.cumsum(1 - sorted_true) # Compute precision at each recall step precision = tp / (tp + fp + K.epsilon()) recall = tp / np.sum(true) # Calculate AP using trapezoidal interpolation ap = np.sum(np.diff(recall) * np.array(precision)[:-1]) ap_scores.append(ap) # Return the mean of all label-specific AP scores return K.mean(K.array(ap_scores))
How to Use These Metrics in Your Model
When compiling your Keras model, pass the custom metrics alongside built-in ones like this:
model.compile( optimizer='adam', loss='binary_crossentropy', # Critical for multi-label tasks! metrics=['accuracy', hamming_loss, multi_label_precision, multi_label_recall, multi_label_f1] )
内容的提问来源于stack exchange,提问作者mouni93

