You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为可变数量标签构建自定义Binary Crossentropy损失函数?

带掩码的批量二元交叉熵自定义损失函数实现

问题场景

给定批量的2D张量,其中y_true中的-1为掩码值,仅需计算非掩码部分的二元交叉熵:

import numpy as np

y_true = np.array([[0., 1., -1., -1.],
                   [1., 0., 1., -1.]])

y_pred = np.array([[0.3, 0.5, 0.2, 0.1],
                   [0.4, 0.3, 0.7, 0.5]])

需要计算有效标签[[0., 1.], [1., 0., 1.]]与对应预测值[[0.3, 0.5], [0.4, 0.3, 0.7]]的二元交叉熵,现有实现仅支持1D张量,需适配2D批量场景。

解决方案:通用批量掩码损失函数

利用tf.boolean_mask过滤掩码元素,适配任意维度的批量张量:

import tensorflow as tf

def masked_binary_crossentropy(y_true, y_pred):
    # 生成掩码:标记y_true中不等于-1的有效元素
    mask = tf.math.not_equal(y_true, tf.constant(-1.0))
    # 过滤掩码元素,得到所有有效标签和预测值的一维张量
    y_true_masked = tf.boolean_mask(y_true, mask)
    y_pred_masked = tf.boolean_mask(y_pred, mask)
    # 计算二元交叉熵
    return tf.keras.losses.binary_crossentropy(y_true_masked, y_pred_masked)

测试验证

# 代入示例数据计算损失
loss = masked_binary_crossentropy(y_true, y_pred)
print(loss.numpy())  # 输出约0.5398

进阶:按样本计算损失再平均

如果需要保留每个样本的损失计算逻辑(避免因单个全掩码样本导致整体错误),可以实现按样本处理的版本:

def per_sample_masked_bce(y_true, y_pred):
    def _single_sample_loss(true, pred):
        mask = tf.math.not_equal(true, -1.0)
        true_masked = tf.boolean_mask(true, mask)
        pred_masked = tf.boolean_mask(pred, mask)
        # 处理无有效标签的样本,返回0避免报错
        if tf.size(true_masked) == 0:
            return 0.0
        return tf.keras.losses.binary_crossentropy(true_masked, pred_masked)
    
    # 对批量中每个样本独立计算损失
    sample_losses = tf.map_fn(_single_sample_loss, (y_true, y_pred), dtype=tf.float32)
    # 返回批量平均损失
    return tf.reduce_mean(sample_losses)

关键说明

  • tf.boolean_mask自动处理多维张量的过滤,无需手动计算有效元素数量或切片,适配任意批量维度
  • 若预测值是logits而非概率,需在binary_crossentropy中设置from_logits=True
  • 进阶版本兼容存在全掩码样本的场景,保证批量计算的鲁棒性

内容的提问来源于stack exchange,提问作者Mykola Zotko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 19:03:39