You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

医学影像二分类:BCE与CrossEntropy效果差异及概率转换咨询

基于Datscan影像的帕金森病二分类问题

我是一名热爱机器学习(ML)的住院医师,希望将ML应用于医学影像领域。我们有一项名为**Datscan闪烁扫描(Datscan scintigraphy)**的检查,通过脑部代谢视图判断患者是否患有帕金森病。该检查耗时约30分钟,老年患者常无法耐受,因此我构建CNN仅用前2个投影(前后位,编号0和60)完成“正常/异常Datscan”二分类,目标输出0-1的“异常Datscan概率”,以便调整阈值控制灵敏度或特异性。

我构建了含887个Datscan的数据集,转换为npy数组(每个含120个128×128像素矩阵),仅使用其中2个,灰度图单通道。尝试PyTorch的VGG类架构,使用BCEWithLogitsLoss时,批量6个样本的输出张量快速收敛至相同值,训练损失下降缓慢,准确率卡在57%左右;改用CrossEntropyLoss后,准确率提升至79%,但需将输出转换为异常概率。

使用BCEWithLogitsLoss的模型代码

class ReseauConvolutionSigmo(nn.Module):
    def __init__(self):
        super(ReseauConvolutionSigmo, self).__init__()
        self.conv1a = nn.Conv2d(2, 64, 3, stride=1)
        self.conv1b = nn.Conv2d(64, 64, 5, stride=1)
        self.pool1 = nn.MaxPool2d(2,2)

        self.conv2a = nn.Conv2d(64, 128, 3, stride=1)
        self.conv2b = nn.Conv2d(128, 128, 3, stride=1)
        self.pool2 = nn.MaxPool2d(2,2)
        
        self.conv3a = nn.Conv2d(128, 256, 3, stride=1)
        self.conv3b = nn.Conv2d(256, 256, 3, stride=1)
        self.pool3 = nn.MaxPool2d(2,2)
                
        self.fc1 = nn.Linear(36864, 84)  
        self.fc2 = nn.Linear(84, 1)       
        
    def forward(self, x):
        x=x.float()
        
        x=self.conv1a(x)
        x=F.relu(x)
        x=self.conv1b(x)
        x=F.relu(x)
        x=self.pool1(x)
        
        x=self.conv2a(x)
        x=F.relu(x)
        x=self.conv2b(x)
        x=F.relu(x)
        x=self.pool2(x)
        
        x=self.conv3a(x)
        x=F.relu(x)
        x=self.conv3b(x)
        x=F.relu(x)
        x=self.pool3(x)
        
        x = torch.flatten(x, 1)  # Flatten the feature maps
        
        try:
            x = F.relu(self.fc1(x))
        except RuntimeError as e:
            e = str(e)
            if e.endswith("Output size is too small"):
                print("Image size is too small.")
            elif "shapes cannot be multiplied" in e:
                required_shape = e[e.index("x") + 1:].split(" ")[0]
                print(f"Linear layer needs to have size: {required_shape}")
            else:
                print(f"Error other: {e}") 
                
        x = self.fc2(x)

        return x
network = ReseauConvolutionSigmo()
n_epochs = 100
criterion = nn.BCEWithLogitsLoss()
optimizer = optim.Adam(network.parameters(), lr=0.001)

train_losses = [ ]
train_counter = [ ]
test_losses = [ ]
test_accuracy = [ ]

network.to(device)
print('******* Evaluation initiale')
test()
for epoch in range(0, n_epochs):
  print('******* Epoch ',epoch)
  train()
  test()

BCEWithLogitsLoss训练异常输出示例

Evaluation initiale
test loss= 0.7000894740570424
Output tensor([[0.0826],
[0.0827],
[0.0827],
[0.0825],
[0.0827],
[0.0827]])
Predicted tensor([[0.], [0.], [0.],[0.],[0.], [0.]])
Datscan tensor([[0.],[0.], [0.],[1.],[0.],[1.]])
Accuracy in test 57.36434108527132 %
...
Epoch  40
train loss= 0.6704580792440817
test loss= 0.6978785312452982
Output tensor([[-0.6284],
[-0.6284],
[-0.6284],
[-0.6284],
[-0.6284],
[-0.6284]])
Predicted tensor([[0.], [0.], [0.],[0.],[0.], [0.]])
Datscan tensor([[0.],[0.], [0.],[0.],[1.],[1.]])
Accuracy in test 56.97674418604651 %

CrossEntropyLoss训练效果示例

Epoch  47
train loss= 0.0002839015607657339
test loss= 1.7087745488627646
correct 203
total 258
Accuracy in test 78.68217054263566 %
Sortie du réseau :
tensor([[-13.9290,   4.8103],
[3.7896,  -9.5477],
[ -3.8057,  -0.1662],
[1.8018,  -3.5083],
[ -3.6199,  -2.2624],
[6.0148, -12.3137]])
Datscan :    tensor([1, 1, 0, 0, 1, 0])
Prédiction :  tensor([1, 0, 1, 0, 1, 0])
...
Epoch  49
train loss= 0.00022697989463199136
test loss= 1.6694575882692397
correct 204
total 258
Accuracy in test 79.06976744186046 %

咨询问题

  • 明明是二分类问题,为何BCEWithLogitsLoss训练效果极差,且批量样本输出快速趋同?调整学习率无改善,是否与优化器有关?
  • 若BCEWithLogitsLoss模型准确率达标,是否可将输出经Sigmoid转换为0-1的异常Datscan概率?
  • CrossEntropyLoss的输出为2×6张量(对应正常/异常类置信度),如何将其转换为异常Datscan概率?

问题解答

  1. BCEWithLogitsLoss效果差的原因

    • 类别不平衡:如果数据集某类占比过高,模型会倾向于输出多数类的预测值,导致批量输出趋同,先检查数据集的类别分布。
    • 数据未归一化:BCEWithLogitsLoss对输入数据的尺度敏感,若影像数据未做标准化(比如归一化到0-1或均值为0方差为1),会导致模型训练不稳定。
    • 模型初始化与结构:单神经元输出的权重初始化不当,加上ReLU的饱和特性,可能让所有样本的输出被拉到同一值;另外检查train()函数是否正确执行梯度清零、反向传播和优化步骤,避免数据加载错误。
    • 优化器调整:Adam不一定是问题,但可以尝试添加权重衰减(weight decay),或换成SGD带动量,打破收敛僵局。
  2. BCEWithLogitsLoss输出转概率
    完全可以。直接对模型输出的logits应用torch.sigmoid(),就能得到0-1之间的异常概率值,可直接用于调整阈值控制灵敏度和特异性。

  3. CrossEntropyLoss输出转异常概率
    对输出的logits应用torch.softmax(dim=1)得到各类别的概率分布,再提取异常类对应的列即可。假设输出张量形状为(6,2),第二列对应异常类,代码示例:

    logits = model_output  # shape (6,2)
    probabilities = torch.softmax(logits, dim=1)
    abnormal_prob = probabilities[:, 1]  # 索引1对应异常类
    

    得到的abnormal_prob就是每个样本的异常概率,范围0-1,可用于阈值调整。


内容的提问来源于stack exchange,提问作者Quazality

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 03:24:57