You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于非离散标签下二元交叉熵损失(BCE)的技术问询

Understanding Binary Cross-Entropy (BCE) with Continuous Targets

Great question! This trips up a lot of folks because most beginner-focused explanations stick to discrete 0/1 labels, but binary cross-entropy (BCE) is way more flexible than that—it works seamlessly with continuous target values between 0 and 1, like the normalized pixel values in your MNIST reconstruction example.

Let’s break this down step by step:

First, Recap: BCE with Discrete 0/1 Labels

The function you shared is the simplified version for discrete targets:

def CrossEntropy(yHat, y):
    if y == 1:
        return -log(yHat)
    else:
        return -log(1 - yHat)

This is just a special case of the general BCE formula, which is:
$$\text{BCE}(y, \hat{y}) = - \left[ y \cdot \log(\hat{y}) + (1 - y) \cdot \log(1 - \hat{y}) \right]$$
When $y$ is exactly 0 or 1, one of the two terms drops out, giving you the simplified function above.

BCE with Continuous Targets (Like MNIST Reconstruction)

When your target $y$ is a continuous value between 0 and 1 (e.g., a normalized MNIST pixel value of 0.7), we interpret this differently: instead of treating $y$ as a hard label, we treat it as the probability parameter of a Bernoulli distribution for that pixel.

For example, a pixel with a true value of 0.7 means: "this pixel's 'true' distribution is a Bernoulli trial where the probability of being 'on' (bright) is 0.7". The BCE loss then measures how well your model's predicted probability $\hat{y}$ matches this true parameter.

The best part? You use the exact same general BCE formula as before. There’s no modification needed—just plug in the continuous $y$ value directly.

Example: MNIST Reconstruction Loss

Let’s say you’ve normalized MNIST’s 0-255 pixel values to the range [0,1]. Your model outputs a prediction $\hat{y}$ for each pixel (usually passed through a sigmoid activation to ensure it’s also in [0,1]). For each pixel, you calculate:
$$- \left[ y_{\text{pixel}} \cdot \log(\hat{y}{\text{pixel}}) + (1 - y{\text{pixel}}) \cdot \log(1 - \hat{y}_{\text{pixel}}) \right]$$
Then you average or sum this value across all pixels in the image (and across your batch) to get the total loss.

Code Example (PyTorch)

Most deep learning frameworks (like PyTorch or TensorFlow) already support continuous targets in their BCE loss implementations. Here’s how it works in PyTorch:

import torch
import torch.nn as nn

# Simulate a batch of normalized MNIST images (32 samples, 28x28=784 pixels)
true_pixels = torch.rand(32, 784)  # Values between 0 and 1
# Model output (after sigmoid to ensure 0-1 range)
predicted_pixels = torch.sigmoid(torch.randn(32, 784))

# Calculate BCE loss—works with continuous targets out of the box!
bce_loss_fn = nn.BCELoss()
total_loss = bce_loss_fn(predicted_pixels, true_pixels)

print(f"Total BCE loss: {total_loss.item():.4f}")

Under the hood, this computes the general BCE formula for every pixel, then averages the results (you can adjust the reduction method with the reduction parameter if needed).

Why This Makes Sense

This approach is also called the Bernoulli negative log-likelihood loss. By treating each continuous target as a Bernoulli parameter, you’re asking your model to predict the probability distribution that best explains the observed pixel value—rather than trying to predict the exact pixel value (which is what MSE loss would do). For image reconstruction tasks, this often leads to sharper outputs compared to MSE, especially for binary or grayscale images.

内容的提问来源于stack exchange,提问作者Matt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:11:52