You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch CNN模型构建报错求助:维度不匹配问题解析

PyTorch单卷积层CNN报错及疑问解答

问题描述

尝试从零构建PyTorch CNN模型,仅使用一层卷积层时出现报错:

RuntimeError: mat1 and mat2 shapes cannot be multiplied (64x197136 and 2704x2704)

不确定单卷积层是否可行,同时不理解out.view(out.size(0),-1)的作用,相关代码如下:

# Craete a neural network from pytorch
class Cnn(nn.Module):
    def __init__(self):
        super(Cnn,self).__init__()
        
        self.layer1 = nn.Sequential(
            nn.Conv2d(in_channels=1, out_channels=16, kernel_size=3, padding=0, stride=2),
            nn.BatchNorm2d(num_features=16),
            nn.ReLU(),
            nn.MaxPool2d(kernel_size=2)
        )

        
        self.fc1 = nn.Linear(3*3*16,10)
        self.dropout = nn.Dropout(0.5)
        self.fc2 = nn.Linear(10,2)
        self.relu = nn.ReLU()
        
        
    def forward(self,x):
        out = self.layer1(x)
        out = out.view(out.size(0),-1)
        out = self.fc1(out)
        out = self.relu(out)
        out = self.fc2(out)
        return out

1. 报错根源

报错核心是全连接层输入维度不匹配。你给fc1设定的输入维度是3*3*16,但经过卷积池化后,特征图展平后的实际维度是197136(对应batch_size=64时的64x197136),两者无法进行矩阵乘法运算。

以常见的224x224单通道输入为例,特征图尺寸计算过程:

  • 卷积层输出尺寸:[(224 - 3 + 0)//2] +1 = 111 → 111x111x16
  • MaxPool2d(kernel_size=2)后尺寸:111//2 = 55 → 55x55x16
  • 展平后总特征数:55*55*16 = 48400,和你设定的3*3*16=144完全不符,这就是报错原因。

2. 单卷积层是否可行?

完全可行,但有两点要注意:

  • 单卷积层的特征提取能力有限,针对猫狗分类这类复杂任务,最终模型精度会偏低;
  • 只要保证卷积后展平的维度和全连接层输入维度一致,模型就能正常运行。

3. out.view(out.size(0),-1)的作用

这行代码是将多维特征图转换为全连接层要求的二维张量:

  • out.size(0)代表batch_size(比如64),保留第一维度不变;
  • -1是让PyTorch自动计算剩余维度的总元素数,把特征图的空间维度(比如55x55)和通道维度(16)合并成一个维度,最终得到[batch_size, 总特征数]的格式——这是全连接层唯一能处理的输入格式。

4. 修正后的代码示例

方式一:手动计算正确维度

针对224x224单通道输入,修正fc1的输入参数:

class Cnn(nn.Module):
    def __init__(self):
        super(Cnn,self).__init__()
        
        self.layer1 = nn.Sequential(
            nn.Conv2d(in_channels=1, out_channels=16, kernel_size=3, padding=0, stride=2),
            nn.BatchNorm2d(16),
            nn.ReLU(),
            nn.MaxPool2d(kernel_size=2)
        )

        # 手动计算后的正确维度:55*55*16=48400
        self.fc1 = nn.Linear(48400,10)
        self.dropout = nn.Dropout(0.5)
        self.fc2 = nn.Linear(10,2)
        self.relu = nn.ReLU()
        
    def forward(self,x):
        out = self.layer1(x)
        out = out.view(out.size(0),-1)
        out = self.fc1(out)
        out = self.relu(out)
        out = self.dropout(out)  # 建议在激活后加dropout,缓解过拟合
        out = self.fc2(out)
        return out

方式二:自适应池化固定输出(更灵活)

用自适应池化固定特征图尺寸,避免手动计算的麻烦:

class Cnn(nn.Module):
    def __init__(self):
        super(Cnn,self).__init__()
        
        self.layer1 = nn.Sequential(
            nn.Conv2d(in_channels=1, out_channels=16, kernel_size=3, padding=0, stride=2),
            nn.BatchNorm2d(16),
            nn.ReLU(),
            # 自适应池化固定输出为3x3,展平后维度固定为3*3*16
            nn.AdaptiveMaxPool2d((3,3))
        )

        self.fc1 = nn.Linear(3*3*16,10)
        self.dropout = nn.Dropout(0.5)
        self.fc2 = nn.Linear(10,2)
        self.relu = nn.ReLU()
        
    def forward(self,x):
        out = self.layer1(x)
        out = out.view(out.size(0),-1)
        out = self.fc1(out)
        out = self.relu(out)
        out = self.dropout(out)
        out = self.fc2(out)
        return out

内容的提问来源于stack exchange,提问作者Auston Barboza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 07:15:33