You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CNN转Linear层维度不匹配:动态生成模型layer_count>9时失败

问题:CNN模型动态生成时Linear层维度计算错误(layer_count>10时失败)

问题描述

针对Fashion MNIST数据集动态生成不同层数的CNN模型,create_cnn_model函数在layer_count为1-9时正常运行,但超过10层就会因最后Linear层维度计算错误报错。模型逻辑为:每3个卷积层后添加MaxPool2D层,完成指定数量的CNN层后添加Flatten和Linear层。

存疑点:

  • create_cnn_model中factor变量的使用是否正确
  • 最后Linear层的维度计算是否准确

相关代码

#Whats the width and height of our images?
W, H = 28, 28 #
#How many values are in the input? We use this to help determine the size of subsequent layers
D = 28*28 #28 * 28 images 
#Hidden layer size
n = 256 
#How many channels are in the input?
C = 1
#how many filters per convolutional layer
n_filters = 32
#How many classes are there?
classes = 10

leak_rate = 0.01

loss_func = nn.CrossEntropyLoss()

# function altered to support optional batch_norm
def cnnLayer(in_filters, out_filters=None, kernel_size=3, batch_norm=False):
    """
    in_filters: how many channels are coming into the layer
    out_filters: how many channels this layer should learn / output, or `None` if we want to have the same number of channels as the input.
    kernel_size: how large the kernel should be
    batch_norm: defines if a batch norm layer should be included
    """
    if out_filters is None:
        out_filters = in_filters #This is a common pattern, so lets automate it as a default if not asked
    padding=kernel_size//2 #padding to stay the same size
    layers = []
    layers.append(nn.Conv2d(in_filters, out_filters, kernel_size, padding=padding))
    if batch_norm:
        layers.append(nn.BatchNorm2d(out_filters))
    layers.append(nn.LeakyReLU(leak_rate))
    return nn.Sequential( # Combine the layer and activation into a single unit
        *layers
    )


def create_cnn_model(layer_count: int, include_bn_layer: bool):
    layers = []
    factor = 1;
    for layer_index in range(layer_count):
        if layer_index == 0:
            layers.append(cnnLayer(C, n_filters))
        elif layer_index % 3 == 0 and layer_index != 0:
            layers.append(nn.MaxPool2d((2,2)))
            layers.append(cnnLayer(factor*n_filters, 2*factor*n_filters))
            factor = factor * 2
        else:
            layers.append(cnnLayer(factor*n_filters))
    layers.append(nn.Flatten())
    layers.append(nn.Linear((D*n_filters)//(factor), classes))
    return nn.Sequential(*layers)

create_cnn_model(10, False)生成的模型结构

Sequential(
  (0): Sequential(
    (0): Conv2d(1, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (1): LeakyReLU(negative_slope=0.01)
  )
  (1): Sequential(
    (0): Conv2d(32, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (1): LeakyReLU(negative_slope=0.01)
  )
  (2): Sequential(
    (0): Conv2d(32, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (1): LeakyReLU(negative_slope=0.01)
  )
  (3): MaxPool2d(kernel_size=(2, 2), stride=(2, 2), padding=0, dilation=1, ceil_mode=False)
  (4): Sequential(
    (0): Conv2d(32, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (1): LeakyReLU(negative_slope=0.01)
  )
  (5): Sequential(
    (0): Conv2d(64, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (1): LeakyReLU(negative_slope=0.01)
  )
  (6): Sequential(
    (0): Conv2d(64, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (1): LeakyReLU(negative_slope=0.01)
  )
  (7): MaxPool2d(kernel_size=(2, 2), stride=(2, 2), padding=0, dilation=1, ceil_mode=False)
  (8): Sequential(
    (0): Conv2d(64, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (1): LeakyReLU(negative_slope=0.01)
  )
  (9): Sequential(
    (0): Conv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (1): LeakyReLU(negative_slope=0.01)
  )
  (10): Sequential(
    (0): Conv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (1): LeakyReLU(negative_slope=0.01)
  )
  (11): MaxPool2d(kernel_size=(2, 2), stride=(2, 2), padding=0, dilation=1, ceil_mode=False)
  (12): Sequential(
    (0): Conv2d(128, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
    (1): LeakyReLU(negative_slope=0.01)
  )
  (13): Flatten(start_dim=1, end_dim=-1)
  (14): Linear(in_features=3136, out_features=10, bias=True)
)

报错信息

RuntimeError: mat1 and mat2 shapes cannot be multiplied (128x2304 and 3136x10)

问题分析

从报错可以看出:Flatten后的特征维度为2304,但Linear层的输入维度被设置为3136,两者不匹配,导致矩阵乘法失败。

1. Linear层维度计算错误

原代码中Linear层的计算逻辑(D*n_filters)//(factor)存在两处问题:

  • 假设每次MaxPool后图像尺寸严格除以2,但当尺寸为奇数时(如7→3),实际是向下取整,无法用简单的除法计算
  • 未考虑最终通道数是factor*n_filters,而非初始的n_filters

2. factor变量逻辑验证

factor的更新逻辑是对的:每触发一次MaxPool(每3个卷积层后),通道数翻倍,factor同步翻倍。但layer_count=10时,会触发3次MaxPool,factor最终变为8,此时通道数为256,图像尺寸经过3次MaxPool后变为3x3,Flatten后的维度应为3*3*256=2304,和报错中的数值一致。


解决方案

方法1:动态计算特征维度(推荐)

通过dummy tensor前向传播,直接获取Flatten前的特征维度,避免手动计算的误差:

import torch

def create_cnn_model(layer_count: int, include_bn_layer: bool):
    layers = []
    factor = 1
    for layer_index in range(layer_count):
        if layer_index == 0:
            layers.append(cnnLayer(C, n_filters))
        elif layer_index % 3 == 0 and layer_index != 0:
            layers.append(nn.MaxPool2d((2,2)))
            layers.append(cnnLayer(factor*n_filters, 2*factor*n_filters))
            factor *= 2
        else:
            layers.append(cnnLayer(factor*n_filters))
    
    # 先构建特征提取部分
    feature_extractor = nn.Sequential(*layers)
    # 用dummy tensor计算输出维度
    dummy_input = torch.randn(1, C, W, H)
    with torch.no_grad():
        feature_output = feature_extractor(dummy_input)
    flatten_dim = feature_output.numel()
    
    # 添加Flatten和Linear层
    layers.append(nn.Flatten())
    layers.append(nn.Linear(flatten_dim, classes))
    return nn.Sequential(*layers)

方法2:修正手动计算逻辑

如果不想使用dummy tensor,可通过公式准确计算每次MaxPool后的图像尺寸:

def calculate_feature_size(w, h, pool_count):
    # MaxPool2d(kernel=2, stride=2, padding=0)的输出尺寸公式
    for _ in range(pool_count):
        w = (w - 2) // 2 + 1
        h = (h - 2) // 2 + 1
    return w, h

def create_cnn_model(layer_count: int, include_bn_layer: bool):
    layers = []
    factor = 1
    for layer_index in range(layer_count):
        if layer_index == 0:
            layers.append(cnnLayer(C, n_filters))
        elif layer_index % 3 == 0 and layer_index != 0:
            layers.append(nn.MaxPool2d((2,2)))
            layers.append(cnnLayer(factor*n_filters, 2*factor*n_filters))
            factor *= 2
        else:
            layers.append(cnnLayer(factor*n_filters))
    
    # 计算MaxPool次数
    pool_count = (layer_count - 1) // 3
    # 计算最终特征图尺寸
    final_w, final_h = calculate_feature_size(W, H, pool_count)
    # 计算Flatten后的维度
    flatten_dim = final_w * final_h * factor * n_filters
    
    layers.append(nn.Flatten())
    layers.append(nn.Linear(flatten_dim, classes))
    return nn.Sequential(*layers)

内容的提问来源于stack exchange,提问作者user2407956

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 21:39:53