PyTorch CNN模型构建报错求助:维度不匹配问题解析
PyTorch单卷积层CNN报错及疑问解答
问题描述
尝试从零构建PyTorch CNN模型,仅使用一层卷积层时出现报错:
RuntimeError: mat1 and mat2 shapes cannot be multiplied (64x197136 and 2704x2704)
不确定单卷积层是否可行,同时不理解out.view(out.size(0),-1)的作用,相关代码如下:
# Craete a neural network from pytorch class Cnn(nn.Module): def __init__(self): super(Cnn,self).__init__() self.layer1 = nn.Sequential( nn.Conv2d(in_channels=1, out_channels=16, kernel_size=3, padding=0, stride=2), nn.BatchNorm2d(num_features=16), nn.ReLU(), nn.MaxPool2d(kernel_size=2) ) self.fc1 = nn.Linear(3*3*16,10) self.dropout = nn.Dropout(0.5) self.fc2 = nn.Linear(10,2) self.relu = nn.ReLU() def forward(self,x): out = self.layer1(x) out = out.view(out.size(0),-1) out = self.fc1(out) out = self.relu(out) out = self.fc2(out) return out
1. 报错根源
报错核心是全连接层输入维度不匹配。你给fc1设定的输入维度是3*3*16,但经过卷积池化后,特征图展平后的实际维度是197136(对应batch_size=64时的64x197136),两者无法进行矩阵乘法运算。
以常见的224x224单通道输入为例,特征图尺寸计算过程:
- 卷积层输出尺寸:
[(224 - 3 + 0)//2] +1 = 111→ 111x111x16 - MaxPool2d(kernel_size=2)后尺寸:
111//2 = 55→ 55x55x16 - 展平后总特征数:
55*55*16 = 48400,和你设定的3*3*16=144完全不符,这就是报错原因。
2. 单卷积层是否可行?
完全可行,但有两点要注意:
- 单卷积层的特征提取能力有限,针对猫狗分类这类复杂任务,最终模型精度会偏低;
- 只要保证卷积后展平的维度和全连接层输入维度一致,模型就能正常运行。
3. out.view(out.size(0),-1)的作用
这行代码是将多维特征图转换为全连接层要求的二维张量:
out.size(0)代表batch_size(比如64),保留第一维度不变;-1是让PyTorch自动计算剩余维度的总元素数,把特征图的空间维度(比如55x55)和通道维度(16)合并成一个维度,最终得到[batch_size, 总特征数]的格式——这是全连接层唯一能处理的输入格式。
4. 修正后的代码示例
方式一:手动计算正确维度
针对224x224单通道输入,修正fc1的输入参数:
class Cnn(nn.Module): def __init__(self): super(Cnn,self).__init__() self.layer1 = nn.Sequential( nn.Conv2d(in_channels=1, out_channels=16, kernel_size=3, padding=0, stride=2), nn.BatchNorm2d(16), nn.ReLU(), nn.MaxPool2d(kernel_size=2) ) # 手动计算后的正确维度:55*55*16=48400 self.fc1 = nn.Linear(48400,10) self.dropout = nn.Dropout(0.5) self.fc2 = nn.Linear(10,2) self.relu = nn.ReLU() def forward(self,x): out = self.layer1(x) out = out.view(out.size(0),-1) out = self.fc1(out) out = self.relu(out) out = self.dropout(out) # 建议在激活后加dropout,缓解过拟合 out = self.fc2(out) return out
方式二:自适应池化固定输出(更灵活)
用自适应池化固定特征图尺寸,避免手动计算的麻烦:
class Cnn(nn.Module): def __init__(self): super(Cnn,self).__init__() self.layer1 = nn.Sequential( nn.Conv2d(in_channels=1, out_channels=16, kernel_size=3, padding=0, stride=2), nn.BatchNorm2d(16), nn.ReLU(), # 自适应池化固定输出为3x3,展平后维度固定为3*3*16 nn.AdaptiveMaxPool2d((3,3)) ) self.fc1 = nn.Linear(3*3*16,10) self.dropout = nn.Dropout(0.5) self.fc2 = nn.Linear(10,2) self.relu = nn.ReLU() def forward(self,x): out = self.layer1(x) out = out.view(out.size(0),-1) out = self.fc1(out) out = self.relu(out) out = self.dropout(out) out = self.fc2(out) return out
内容的提问来源于stack exchange,提问作者Auston Barboza
相关产品推荐
相关产品推荐

