You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CNN训练异常:Loss停滞于1且输出等概率,求排查方案

问题分析:CNN4分类任务Loss停滞、输出等概率的原因

问题背景

首次训练CNN完成4分类图像任务,数据集含4类,每类900张图片。训练105轮后,Loss始终停滞在1左右且曲线有尖峰,推理时模型输出四类概率均约25%(等概率随机猜测)。

模型代码

class RedClasificacion(nn.Module):
def __init__(self, *args, **kwargs) -> None:
    super().__init__(*args, **kwargs)
    self.conv1 = nn.Conv2d(3, 32, kernel_size=3, stride=3) # 64 canales, 36, 36
    self.norm1 = nn.BatchNorm2d(32)
    self.act1 = nn.ReLU()
    self.conv2 = nn.Conv2d(32, 64, kernel_size=3, stride=3) # 128 canales, 12, 12
    self.norm2 = nn.BatchNorm2d(64)
    self.drop1 = nn.Dropout(0.35)
    self.act2 = nn.ReLU()
    self.pool1 = nn.MaxPool2d(kernel_size=2, stride=2) # 64 canales, 6, 6
    self.conv3 = nn.Conv2d(64, 128, kernel_size=2, stride=2) # 256 canales, 6, 6
    self.norm3 = nn.BatchNorm2d(128)
    self.drop2 = nn.Dropout(0.35)
    self.act3 = nn.ReLU()
    self.pool2 = nn.MaxPool2d(kernel_size=2, stride=2) # 32 canales, 6, 6
    self.flat = nn.Flatten(start_dim=1, end_dim=-1)
    self.lin1 = nn.Linear(32 * 4 * 4, 128)
    self.norm4 = nn.BatchNorm1d(128)
    self.drop3 = nn.Dropout(0.35)
    self.act4 = nn.ReLU()
    self.lin3 = nn.Linear(128, 4)
    self.softf = nn.Softmax(dim=1)

数据增强代码

transform = transforms.Compose([
    transforms.RandomHorizontalFlip(0.5),
    transforms.RandomRotation(180),
    transforms.RandomInvert(),
    transforms.Resize(size=(144, 144)),
    transforms.ToTensor(),
    transforms.Normalize(mean=(0.5, 0.5, 0.5), std=(0.5, 0.5, 0.5))])

核心问题原因

  • 模型结构维度完全不匹配(最关键)
    代码注释与实际张量维度计算完全不符,全连接层输入维度存在错误:

    • 输入图像为144x144,经conv1(kernel=3, stride=3)后输出尺寸应为(144-3)/3 +1 = 48(即(32,48,48)),而非注释的36x36;
    • 经conv2(kernel=3, stride=3)后输出尺寸为(48-3)/3 +1=16(即(64,16,16)),而非注释的12x12;
    • pool1(2x2池化)后输出(64,8,8),而非注释的6x6;
    • conv3(kernel=2, stride=2)后输出(8-2)/2 +1=4(即(128,4,4)),而非注释的6x6;
    • pool2(2x2池化)后输出(128,2,2),flatten后维度为128*2*2=512,但代码中lin1的输入写为32*4*4,虽数值巧合相等,但注释完全错误,若中间层维度计算偏差会直接导致模型参数更新混乱,无法学习有效特征。
  • Softmax与损失函数冲突
    若训练时使用CrossEntropyLoss,该损失函数内部已集成LogSoftmax和NLLLoss,模型末尾额外添加Softmax层会导致梯度消失,模型无法有效更新参数,最终退化为随机猜测。

  • 数据增强过度
    RandomRotation(180)和RandomInvert会严重破坏图像的关键特征(如方向、颜色信息),若这些特征是区分四类的核心依据,过度增强会让模型无法学到有效分类特征,导致Loss无法下降。

  • Dropout使用不合理
    多处使用0.35的Dropout且直接加在卷积层之后,结合BatchNorm会过度抑制模型的学习信号,尤其是训练初期,模型还未学到稳定特征就被频繁Dropout干扰,难以收敛。

内容的提问来源于stack exchange,提问作者Javi22

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 12:57:13