CNN训练异常:Loss停滞于1且输出等概率,求排查方案
问题分析:CNN4分类任务Loss停滞、输出等概率的原因
问题背景
首次训练CNN完成4分类图像任务,数据集含4类,每类900张图片。训练105轮后,Loss始终停滞在1左右且曲线有尖峰,推理时模型输出四类概率均约25%(等概率随机猜测)。
模型代码
class RedClasificacion(nn.Module): def __init__(self, *args, **kwargs) -> None: super().__init__(*args, **kwargs) self.conv1 = nn.Conv2d(3, 32, kernel_size=3, stride=3) # 64 canales, 36, 36 self.norm1 = nn.BatchNorm2d(32) self.act1 = nn.ReLU() self.conv2 = nn.Conv2d(32, 64, kernel_size=3, stride=3) # 128 canales, 12, 12 self.norm2 = nn.BatchNorm2d(64) self.drop1 = nn.Dropout(0.35) self.act2 = nn.ReLU() self.pool1 = nn.MaxPool2d(kernel_size=2, stride=2) # 64 canales, 6, 6 self.conv3 = nn.Conv2d(64, 128, kernel_size=2, stride=2) # 256 canales, 6, 6 self.norm3 = nn.BatchNorm2d(128) self.drop2 = nn.Dropout(0.35) self.act3 = nn.ReLU() self.pool2 = nn.MaxPool2d(kernel_size=2, stride=2) # 32 canales, 6, 6 self.flat = nn.Flatten(start_dim=1, end_dim=-1) self.lin1 = nn.Linear(32 * 4 * 4, 128) self.norm4 = nn.BatchNorm1d(128) self.drop3 = nn.Dropout(0.35) self.act4 = nn.ReLU() self.lin3 = nn.Linear(128, 4) self.softf = nn.Softmax(dim=1)
数据增强代码
transform = transforms.Compose([ transforms.RandomHorizontalFlip(0.5), transforms.RandomRotation(180), transforms.RandomInvert(), transforms.Resize(size=(144, 144)), transforms.ToTensor(), transforms.Normalize(mean=(0.5, 0.5, 0.5), std=(0.5, 0.5, 0.5))])
核心问题原因
模型结构维度完全不匹配(最关键)
代码注释与实际张量维度计算完全不符,全连接层输入维度存在错误:- 输入图像为144x144,经
conv1(kernel=3, stride=3)后输出尺寸应为(144-3)/3 +1 = 48(即(32,48,48)),而非注释的36x36; - 经
conv2(kernel=3, stride=3)后输出尺寸为(48-3)/3 +1=16(即(64,16,16)),而非注释的12x12; pool1(2x2池化)后输出(64,8,8),而非注释的6x6;conv3(kernel=2, stride=2)后输出(8-2)/2 +1=4(即(128,4,4)),而非注释的6x6;pool2(2x2池化)后输出(128,2,2),flatten后维度为128*2*2=512,但代码中lin1的输入写为32*4*4,虽数值巧合相等,但注释完全错误,若中间层维度计算偏差会直接导致模型参数更新混乱,无法学习有效特征。
- 输入图像为144x144,经
Softmax与损失函数冲突
若训练时使用CrossEntropyLoss,该损失函数内部已集成LogSoftmax和NLLLoss,模型末尾额外添加Softmax层会导致梯度消失,模型无法有效更新参数,最终退化为随机猜测。数据增强过度
RandomRotation(180)和RandomInvert会严重破坏图像的关键特征(如方向、颜色信息),若这些特征是区分四类的核心依据,过度增强会让模型无法学到有效分类特征,导致Loss无法下降。Dropout使用不合理
多处使用0.35的Dropout且直接加在卷积层之后,结合BatchNorm会过度抑制模型的学习信号,尤其是训练初期,模型还未学到稳定特征就被频繁Dropout干扰,难以收敛。
内容的提问来源于stack exchange,提问作者Javi22
相关产品推荐
相关产品推荐

