You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch简易CNN模型训练损失无变化的优化咨询

图像分类模型训练停滞问题求助

任务概述

约15类图像分类任务,类别仅依据颜色划分,无复杂细节。训练100轮的损失与准确率曲线显示,训练50轮后模型性能进入平台期,无明显提升。

所用模型(SimNet1)

模型结构输出

SimNet1(
  (conv1): Sequential(
    (0): Conv2d(3, 4, kernel_size=(5, 5), stride=(1, 1))
    (1): ReLU()
    (2): MaxPool2d(kernel_size=2, stride=2, padding=0, dilation=1, ceil_mode=False)
    (3): Conv2d(4, 6, kernel_size=(5, 5), stride=(1, 1))
    (4): ReLU()
    (5): MaxPool2d(kernel_size=2, stride=2, padding=0, dilation=1, ceil_mode=False)
  )
  (mlp1): Sequential(
    (0): LazyLinear(in_features=0, out_features=120, bias=True)
    (1): ReLU()
    (2): Linear(in_features=120, out_features=60, bias=True)
    (3): ReLU()
    (4): Linear(in_features=60, out_features=13, bias=True)
  )
)

模型定义代码

class SimNet1(nn.Module):
    def __init__(self, conv_out_1, conv_out_2, hid_dim_1, hid_dim_2, num_classes, kernel_size):
        super().__init__()

        # === Start Conv Layers ===
        self.conv1 = nn.Sequential(
            nn.Conv2d(3, conv_out_1, kernel_size),
            nn.ReLU(),
            nn.MaxPool2d(2, 2),
            nn.Conv2d(conv_out_1, conv_out_2, kernel_size),
            nn.ReLU(),
            nn.MaxPool2d(2, 2)
        )

        # === End Conv Layers ===


        # === Start MLP Layers ===
        self.mlp1 = nn.Sequential(
            nn.LazyLinear(hid_dim_1), # if using nn.Linear(), in_dim determined by final conv_out * (image dim after conv)^2
            nn.ReLU(),
            nn.Linear(hid_dim_1, hid_dim_2),
            nn.ReLU(),
            nn.Linear(hid_dim_2, num_classes) # final output dimension matches num of classes
        )

        # === End MLP Layers ===


    def forward(self, x):
        x = self.conv1(x)

        # Flatten tensor except batch
        x = torch.flatten(x, 1)

        x = self.mlp1(x)

        return x

训练配置

  • 优化器:SGD,学习率0.02、动量0.9
  • 学习率调度器:StepLR,步长7、gamma=0.1

训练现状

训练50余轮后,训练损失始终维持在2.53左右,准确率稳定在0.153;验证集损失与准确率处于相近水平,无明显波动。

此前曾使用修改最后一层输出维度的ResNet18,训练准确率可达0.6,验证准确率约0.5,但存在过拟合问题(训练损失远低于验证损失),因此更换为当前简易模型。

已尝试的优化方案

  • 增加卷积通道数、隐藏层单元数,小幅提升模型复杂度,无改善
  • 调整学习率(从0.0002改为0.02),仍无改善,怀疑权重衰减与学习率不匹配

求助需求

如何解决该简易模型训练损失无提升的问题?


内容的提问来源于stack exchange,提问作者AlgoManiac

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 16:05:42