CNN转Linear层维度不匹配:动态生成模型layer_count>9时失败
问题:CNN模型动态生成时Linear层维度计算错误(layer_count>10时失败)
问题描述
针对Fashion MNIST数据集动态生成不同层数的CNN模型,create_cnn_model函数在layer_count为1-9时正常运行,但超过10层就会因最后Linear层维度计算错误报错。模型逻辑为:每3个卷积层后添加MaxPool2D层,完成指定数量的CNN层后添加Flatten和Linear层。
存疑点:
create_cnn_model中factor变量的使用是否正确- 最后Linear层的维度计算是否准确
相关代码
#Whats the width and height of our images? W, H = 28, 28 # #How many values are in the input? We use this to help determine the size of subsequent layers D = 28*28 #28 * 28 images #Hidden layer size n = 256 #How many channels are in the input? C = 1 #how many filters per convolutional layer n_filters = 32 #How many classes are there? classes = 10 leak_rate = 0.01 loss_func = nn.CrossEntropyLoss() # function altered to support optional batch_norm def cnnLayer(in_filters, out_filters=None, kernel_size=3, batch_norm=False): """ in_filters: how many channels are coming into the layer out_filters: how many channels this layer should learn / output, or `None` if we want to have the same number of channels as the input. kernel_size: how large the kernel should be batch_norm: defines if a batch norm layer should be included """ if out_filters is None: out_filters = in_filters #This is a common pattern, so lets automate it as a default if not asked padding=kernel_size//2 #padding to stay the same size layers = [] layers.append(nn.Conv2d(in_filters, out_filters, kernel_size, padding=padding)) if batch_norm: layers.append(nn.BatchNorm2d(out_filters)) layers.append(nn.LeakyReLU(leak_rate)) return nn.Sequential( # Combine the layer and activation into a single unit *layers ) def create_cnn_model(layer_count: int, include_bn_layer: bool): layers = [] factor = 1; for layer_index in range(layer_count): if layer_index == 0: layers.append(cnnLayer(C, n_filters)) elif layer_index % 3 == 0 and layer_index != 0: layers.append(nn.MaxPool2d((2,2))) layers.append(cnnLayer(factor*n_filters, 2*factor*n_filters)) factor = factor * 2 else: layers.append(cnnLayer(factor*n_filters)) layers.append(nn.Flatten()) layers.append(nn.Linear((D*n_filters)//(factor), classes)) return nn.Sequential(*layers)
create_cnn_model(10, False)生成的模型结构
Sequential( (0): Sequential( (0): Conv2d(1, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (1): LeakyReLU(negative_slope=0.01) ) (1): Sequential( (0): Conv2d(32, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (1): LeakyReLU(negative_slope=0.01) ) (2): Sequential( (0): Conv2d(32, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (1): LeakyReLU(negative_slope=0.01) ) (3): MaxPool2d(kernel_size=(2, 2), stride=(2, 2), padding=0, dilation=1, ceil_mode=False) (4): Sequential( (0): Conv2d(32, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (1): LeakyReLU(negative_slope=0.01) ) (5): Sequential( (0): Conv2d(64, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (1): LeakyReLU(negative_slope=0.01) ) (6): Sequential( (0): Conv2d(64, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (1): LeakyReLU(negative_slope=0.01) ) (7): MaxPool2d(kernel_size=(2, 2), stride=(2, 2), padding=0, dilation=1, ceil_mode=False) (8): Sequential( (0): Conv2d(64, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (1): LeakyReLU(negative_slope=0.01) ) (9): Sequential( (0): Conv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (1): LeakyReLU(negative_slope=0.01) ) (10): Sequential( (0): Conv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (1): LeakyReLU(negative_slope=0.01) ) (11): MaxPool2d(kernel_size=(2, 2), stride=(2, 2), padding=0, dilation=1, ceil_mode=False) (12): Sequential( (0): Conv2d(128, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1)) (1): LeakyReLU(negative_slope=0.01) ) (13): Flatten(start_dim=1, end_dim=-1) (14): Linear(in_features=3136, out_features=10, bias=True) )
报错信息
RuntimeError: mat1 and mat2 shapes cannot be multiplied (128x2304 and 3136x10)
问题分析
从报错可以看出:Flatten后的特征维度为2304,但Linear层的输入维度被设置为3136,两者不匹配,导致矩阵乘法失败。
1. Linear层维度计算错误
原代码中Linear层的计算逻辑(D*n_filters)//(factor)存在两处问题:
- 假设每次MaxPool后图像尺寸严格除以2,但当尺寸为奇数时(如7→3),实际是向下取整,无法用简单的除法计算
- 未考虑最终通道数是
factor*n_filters,而非初始的n_filters
2. factor变量逻辑验证
factor的更新逻辑是对的:每触发一次MaxPool(每3个卷积层后),通道数翻倍,factor同步翻倍。但layer_count=10时,会触发3次MaxPool,factor最终变为8,此时通道数为256,图像尺寸经过3次MaxPool后变为3x3,Flatten后的维度应为3*3*256=2304,和报错中的数值一致。
解决方案
方法1:动态计算特征维度(推荐)
通过dummy tensor前向传播,直接获取Flatten前的特征维度,避免手动计算的误差:
import torch def create_cnn_model(layer_count: int, include_bn_layer: bool): layers = [] factor = 1 for layer_index in range(layer_count): if layer_index == 0: layers.append(cnnLayer(C, n_filters)) elif layer_index % 3 == 0 and layer_index != 0: layers.append(nn.MaxPool2d((2,2))) layers.append(cnnLayer(factor*n_filters, 2*factor*n_filters)) factor *= 2 else: layers.append(cnnLayer(factor*n_filters)) # 先构建特征提取部分 feature_extractor = nn.Sequential(*layers) # 用dummy tensor计算输出维度 dummy_input = torch.randn(1, C, W, H) with torch.no_grad(): feature_output = feature_extractor(dummy_input) flatten_dim = feature_output.numel() # 添加Flatten和Linear层 layers.append(nn.Flatten()) layers.append(nn.Linear(flatten_dim, classes)) return nn.Sequential(*layers)
方法2:修正手动计算逻辑
如果不想使用dummy tensor,可通过公式准确计算每次MaxPool后的图像尺寸:
def calculate_feature_size(w, h, pool_count): # MaxPool2d(kernel=2, stride=2, padding=0)的输出尺寸公式 for _ in range(pool_count): w = (w - 2) // 2 + 1 h = (h - 2) // 2 + 1 return w, h def create_cnn_model(layer_count: int, include_bn_layer: bool): layers = [] factor = 1 for layer_index in range(layer_count): if layer_index == 0: layers.append(cnnLayer(C, n_filters)) elif layer_index % 3 == 0 and layer_index != 0: layers.append(nn.MaxPool2d((2,2))) layers.append(cnnLayer(factor*n_filters, 2*factor*n_filters)) factor *= 2 else: layers.append(cnnLayer(factor*n_filters)) # 计算MaxPool次数 pool_count = (layer_count - 1) // 3 # 计算最终特征图尺寸 final_w, final_h = calculate_feature_size(W, H, pool_count) # 计算Flatten后的维度 flatten_dim = final_w * final_h * factor * n_filters layers.append(nn.Flatten()) layers.append(nn.Linear(flatten_dim, classes)) return nn.Sequential(*layers)
内容的提问来源于stack exchange,提问作者user2407956
相关产品推荐
相关产品推荐

