You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SSD目标检测网络训练中边界框MAE骤降但预测失效问题排查

SSD目标检测模型训练异常问题排查求助

问题描述

我正在搭建SSD目标检测神经网络,当前使用600张700×700的单类别训练样本,计划扩展至1000张或更多。训练过程中,预测边界框坐标的平均绝对误差(MAE)出现异常骤降,但模型对任意图像的预测结果完全错误。我已尝试调整学习率、训练轮数、批量大小,甚至新增卷积层,问题仍未解决,恳请协助排查根源。

训练代码

clsLoss = nn.CrossEntropyLoss(reduction='none')
bboxLoss = torch.nn.L1Loss(reduction='none')

batch_size = 32
train_iter, _ = datasetting.loadData(batch_size)

device, net = try_gpu(), model.TinySSD(num_classes=1)
trainer = torch.optim.SGD(net.parameters(), lr=3e-4, weight_decay=5e-4)
num_epochs, timer = 11, Timer()
animator = Animator(xlabel='epoch', xlim=[1, num_epochs],
                        legend=['class error', 'bbox mae'])
net = net.to(device)
for epoch in range(num_epochs):
    metric = Accumulator(4)
    net.train()
    for features, target in train_iter:
        timer.start()
        trainer.zero_grad()
        X, Y = features.to(device), target.to(device)
        
        anchors, cls_preds, bbox_preds = net(X)
        
        bbox_labels, bbox_masks, cls_labels = anchor.multiboxTarget(anchors, Y)
     
        l = calcLoss(cls_preds, cls_labels, bbox_preds, bbox_labels,
                      bbox_masks)
        l.mean().backward()
        trainer.step()
        metric.add(clsEval(cls_preds, cls_labels), cls_labels.numel(),
                   bboxEval(bbox_preds, bbox_labels, bbox_masks),
                   bbox_labels.numel())
    cls_err, bbox_mae = 1 - metric[0] / metric[1], metric[2] / metric[3]
    animator.add(epoch + 1, (cls_err, bbox_mae))
print(f'class err {cls_err:.2e}, bbox mae {bbox_mae:.2e}')
print(f'{len(train_iter.dataset) / timer.stop():.1f} examples/sec on '
      f'{str(device)}')

网络结构代码

def downSampleBlk(in_channels, out_channels):
    blk = []
    for _ in range(2):
        blk.append(nn.Conv2d(in_channels, out_channels,
                             kernel_size=3, padding=1))
        blk.append(nn.BatchNorm2d(out_channels))
        blk.append(nn.ReLU())
        in_channels = out_channels
    blk.append(nn.MaxPool2d(2))
    return nn.Sequential(*blk)


def base_net():
    blk = []
    num_filters = [3, 16, 32, 64]
    for i in range(len(num_filters) - 1):
        blk.append(downSampleBlk(num_filters[i], num_filters[i+1]))
    return nn.Sequential(*blk)


def getBlk(i):
    if i == 0:
        blk = base_net()
    elif i == 1:
        blk = downSampleBlk(64, 128)
    elif i == 5:
        blk = nn.AdaptiveMaxPool2d((1,1))
    else:
        blk = downSampleBlk(128, 128)
    return blk

损失计算与预测代码

cls_loss = nn.CrossEntropyLoss(reduction='none')
bbox_loss = nn.L1Loss(reduction='none')

def calcLoss(cls_preds, cls_labels, bbox_preds, bbox_labels, bbox_masks):
    batch_size, num_classes = cls_preds.shape[0], cls_preds.shape[2]
    cls = clsLoss(cls_preds.reshape(-1, num_classes),
                   cls_labels.reshape(-1)).reshape(batch_size, -1).mean(dim=1)
    bbox = bboxLoss(bbox_preds * bbox_masks,
                     bbox_labels * bbox_masks).mean(dim=1)
    #print(bbox)
    return cls + bbox

def predict(X):
    net.eval()
    anchors, cls_preds, bbox_preds = net(X.to(device))
    cls_probs = F.softmax(cls_preds, dim=2).permute(0, 2, 1)
    output = anchor.multiboxDetection(cls_probs, bbox_preds, anchors)
    idx = [i for i, row in enumerate(output[0]) if row[0] != -1]
    return output[0, idx]

相关结果

  • 边界框MAE曲线:边界框MAE曲线
  • 训练样本预测结果(目标为红色路标):训练样本预测结果
  • 测试样本预测结果:测试样本预测结果
  • 边界框损失曲线:边界框损失曲线
  • 分类损失曲线:分类损失曲线

内容的提问来源于stack exchange,提问作者KarimCool

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 15:27:52