You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AlexNet训练正常但打印每层输出时报mat1和mat2无法相乘的RuntimeError

问题原因

网络正常训练可运行是因为forward方法中,卷积层conv输出4维特征张量(形状为[batch_size, 256, 5, 5])后,主动做了展平操作feature.view(img.shape[0], -1),把4维张量转成了2维的[batch_size, 6400],刚好匹配全连接层第一个线性层的输入维度要求256*5*5=6400。

但调试代码直接遍历网络的顶层子模块(conv和fc两个Sequential),没有做展平操作:

  • 跑完conv层后输出的X形状是[1, 256, 5, 5],属于4维张量
  • 直接把这个4维张量传入fc层时,PyTorch的nn.Linear默认只对输入的最后一维做全连接计算,会自动把前面的所有维度都当作批量维度,所以实际传入线性层的张量会被识别为形状(1*256*5, 5) = 1280x5,和线性层要求的输入维度6400不匹配,因此抛出形状相乘错误。
解决方案

提供两种可选方案:

方案1:修改调试代码,卷积层输出后加展平操作

无需改动原网络结构,仅调整调试逻辑即可:

X = torch.randn(1,1,224,224)
for name,layer in net.named_children():
    X = layer(X)
    # 卷积层输出后做展平,再输入全连接层
    if name == 'conv':
        X = X.view(X.shape[0], -1)
    print(name, X.shape)

方案2:把展平操作写入网络结构中

直接在conv序列末尾加nn.Flatten()层,后续训练、调试都不需要额外处理形状适配问题,修改后的网络定义如下:

class AlexNet(nn.Module):
    def __init__(self):
      super(AlexNet, self).__init__()
      self.conv = nn.Sequential(
          nn.Conv2d(1, 96, 11, 4),
          nn.ReLU(),
          nn.MaxPool2d(3, 2),
          nn.Conv2d(96, 256, 5, 1, 2),
          nn.ReLU(),
          nn.MaxPool2d(3, 2),
          nn.Conv2d(256, 384, 3, 1, 1),
          nn.ReLU(),
          nn.Conv2d(384, 384, 3, 1, 1),
          nn.ReLU(),
          nn.Conv2d(384, 256, 3, 1, 1),
          nn.ReLU(),
          nn.MaxPool2d(3, 2),
          # 新增展平层,输出维度为[batch_size, 6400]
          nn.Flatten()
         )
      self.fc = nn.Sequential(
          nn.Linear(256*5*5, 4096),
          nn.ReLU(),
          nn.Dropout(0.5),
          nn.Linear(4096, 4096),
          nn.ReLU(),
          nn.Dropout(0.5),
          nn.Linear(4096, 10)
        )
      
    def forward(self, img):
      feature = self.conv(img)
      output = self.fc(feature)
      return output

修改后原有调试代码无需改动即可正常运行,逐层输出特征形状。


内容的提问来源于stack exchange,提问作者machine no learning

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 13:06:11