You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Keras的多标签图像分类自定义指标技术咨询

Got it, let's break down how to build custom Keras metrics for your 39-label multi-image classification task. Since each image can have multiple features, standard single-class metrics won’t cut it—we need metrics tailored to multi-label scenarios. Below are practical, ready-to-use custom metrics that align with your data setup (images shaped (1814,204,204,3) and labels shaped (1814,39)).

Keras Custom Metrics for 39-Label Multi-Image Classification

First, a quick reminder: for multi-label tasks, always use binary_crossentropy as your loss function (each label is an independent binary classification problem). Now let's dive into the metrics:


1. Hamming Loss (Lower = Better)

This metric measures the fraction of misclassified labels across all samples and labels—it’s perfect for getting a holistic view of how often your model gets any label wrong.

import tensorflow as tf
from tensorflow.keras import backend as K

def hamming_loss(y_true, y_pred):
    # Threshold predictions to binary 0/1 using 0.5 as the cutoff
    y_pred_bin = K.round(K.clip(y_pred, 0, 1))
    # Count mismatched labels (false positives + false negatives)
    mismatches = K.sum(K.abs(y_true - y_pred_bin))
    # Total number of label positions across all samples
    total_labels = K.cast(K.prod(K.shape(y_true)), dtype=K.floatx())
    return mismatches / total_labels

2. Multi-Label Precision (Higher = Better)

Precision tells you how many of the labels your model predicted as positive are actually positive—great if you want to minimize false alarms (e.g., not flagging a feature that isn’t present).

def multi_label_precision(y_true, y_pred):
    y_pred_bin = K.round(K.clip(y_pred, 0, 1))
    # True positives: labels predicted positive that are actually positive
    true_positives = K.sum(K.round(y_true * y_pred_bin))
    # Total predicted positives: all labels the model flagged as positive
    predicted_positives = K.sum(y_pred_bin)
    # Avoid division by zero if no positives are predicted in a batch
    return K.switch(K.equal(predicted_positives, 0), 0.0, true_positives / predicted_positives)

3. Multi-Label Recall (Higher = Better)

Recall measures how many of the actual positive labels your model correctly identified—critical if you don’t want to miss important features (e.g., failing to detect a key feature that’s present).

def multi_label_recall(y_true, y_pred):
    y_pred_bin = K.round(K.clip(y_pred, 0, 1))
    true_positives = K.sum(K.round(y_true * y_pred_bin))
    # Total actual positives: all labels that are truly present in the data
    actual_positives = K.sum(y_true)
    # Avoid division by zero if no true positives exist in a batch
    return K.switch(K.equal(actual_positives, 0), 0.0, true_positives / actual_positives)

4. Multi-Label F1-Score (Higher = Better)

F1 is the harmonic mean of precision and recall—it balances both metrics, making it ideal when you care about avoiding both false positives and missed true positives.

def multi_label_f1(y_true, y_pred):
    precision = multi_label_precision(y_true, y_pred)
    recall = multi_label_recall(y_true, y_pred)
    # Calculate F1: 2*(P*R)/(P+R)
    return K.switch(K.equal(precision + recall, 0), 0.0, 2 * ((precision * recall) / (precision + recall)))

5. Mean Average Precision (mAP) (Higher = Better)

mAP is a more advanced metric that computes precision for each label at different recall thresholds, then averages those scores. It gives you a detailed view of performance across all 39 labels.

import numpy as np

def mean_average_precision(y_true, y_pred):
    # Convert tensors to numpy arrays (works with TensorFlow 2.x eager execution)
    y_true_np = K.eval(y_true)
    y_pred_np = K.eval(y_pred)
    
    ap_scores = []
    for label_idx in range(y_true_np.shape[1]):
        # Extract true and predicted values for the current label
        true = y_true_np[:, label_idx]
        pred = y_pred_np[:, label_idx]
        
        # Sort predictions in descending order of confidence
        sorted_indices = np.argsort(pred)[::-1]
        sorted_true = true[sorted_indices]
        
        # Calculate cumulative true positives and false positives
        tp = np.cumsum(sorted_true)
        fp = np.cumsum(1 - sorted_true)
        
        # Compute precision at each recall step
        precision = tp / (tp + fp + K.epsilon())
        recall = tp / np.sum(true)
        
        # Calculate AP using trapezoidal interpolation
        ap = np.sum(np.diff(recall) * np.array(precision)[:-1])
        ap_scores.append(ap)
    
    # Return the mean of all label-specific AP scores
    return K.mean(K.array(ap_scores))

How to Use These Metrics in Your Model

When compiling your Keras model, pass the custom metrics alongside built-in ones like this:

model.compile(
    optimizer='adam',
    loss='binary_crossentropy',  # Critical for multi-label tasks!
    metrics=['accuracy', hamming_loss, multi_label_precision, multi_label_recall, multi_label_f1]
)

内容的提问来源于stack exchange,提问作者mouni93

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:12:00