You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch CNN全连接层形状不匹配:mat1与mat2无法相乘报错

问题描述
  • 数据集拆分后形状:X_train (98, 1, 40, 844)、X_val (21, 1, 40, 844)、X_test (21, 1, 40, 844)
  • CNN模型训练过程正常,但在验证集执行模型解释时,forward函数中x = F.relu(self.fc1(x))处触发错误:
    RuntimeError: mat1 and mat2 shapes cannot be multiplied (32x2110 and 67520x128)
    
  • 已尝试修改前向传播函数、调整层形状,问题仍未解决

用户提供的代码如下:

from fastai.vision.all import *
import librosa
import numpy as np
from sklearn.model_selection import train_test_split
import torch
import torch.nn as nn
from torchsummary import summary

[...] #labels in y can be [0,1,2,3]

# Split the data
X_train, X_temp, y_train, y_temp = train_test_split(X, y, test_size=0.3, random_state=42)
X_val, X_test, y_val, y_test = train_test_split(X_temp, y_temp, test_size=0.5, random_state=42)

# Reshape data for CNN input (add channel dimension)
X_train = X_train[:, np.newaxis, :, :]
X_val = X_val[:, np.newaxis, :, :]
X_test = X_test[:, np.newaxis, :, :]

#X_train.shape, X_val.shape, X_test.shape
#((98, 1, 40, 844), (21, 1, 40, 844), (21, 1, 40, 844))

class DraftCNN(nn.Module):
    def __init__(self):
        super(DraftCNN, self).__init__()
        self.conv1 = nn.Conv2d(1, 16, kernel_size=3, stride=1, padding=1)
        self.pool = nn.MaxPool2d(kernel_size=2, stride=2, padding=0)
        self.conv2 = nn.Conv2d(16, 32, kernel_size=3, stride=1, padding=1)
        
        # Calculate flattened size based on input dimensions
        with torch.no_grad():
            dummy_input = torch.zeros(1, 1, 40, 844)  # shape of one input sample
            dummy_output = self.pool(self.conv2(self.pool(F.relu(self.conv1(dummy_input)))))
            self.flattened_size = dummy_output.view(dummy_output.size(0), -1).size(1)
        
        self.fc1 = nn.Linear(self.flattened_size, 128)
        self.fc2 = nn.Linear(128, 4)

    def forward(self, x):
        x = self.pool(F.relu(self.conv1(x)))
        x = self.pool(F.relu(self.conv2(x)))
        x = x.view(x.size(0), -1)  # Flatten the output of convolutions
        x = F.relu(self.fc1(x))
        x = self.fc2(x)
        return x


# Initialize the model and the Learner
model = AudioCNN()
learn = Learner(dls, model, loss_func=CrossEntropyLossFlat(), metrics=[accuracy, Precision(average='macro'),  Recall(average='macro'), F1Score(average='macro')])

# Train the model
learn.fit_one_cycle(8)

print(summary(model, (1, 40, 844)))

# Create a DataLoader for the validation set
valid_dl = learn.dls.test_dl(X_val, y_val)

# Get predictions and interpret them on the validation set
interp = ClassificationInterpretation.from_learner(learn, dl=valid_dl)
interp.plot_confusion_matrix()
interp.plot_top_losses(5)
错误原因分析

报错核心是全连接层输入维度不匹配:模型初始化时计算的扁平化维度(67520)和实际推理时输入经过卷积池化后的扁平化维度(2110)不一致,导致矩阵乘法无法执行。

具体原因拆解:

  1. 代码定义的模型类是DraftCNN,但实例化时用了AudioCNN(),如果AudioCNN是未定义类或结构与DraftCNN不同,会直接导致全连接层维度计算错误
  2. 即使AudioCNN是DraftCNN的别名,也可能因为test_dl的输入形状、预处理与训练集不一致,导致卷积池化后的输出维度偏离预期
解决方案

1. 修正模型实例化错误

将模型实例化代码从AudioCNN()改为DraftCNN(),确保使用的是你定义的正确模型结构:

# 错误代码
model = AudioCNN()
# 修正后
model = DraftCNN()

2. 验证卷积池化后的维度正确性

手动计算输入经过卷积池化后的维度,确认与代码自动计算结果一致:
输入形状(1,1,40,844)的处理流程:

  • conv1输出:(1,16,40,844)(卷积核3×3,padding=1,维度不变)
  • 第一次pool输出:(1,16,20,422)(池化核2×2,步长2,维度减半)
  • conv2输出:(1,32,20,422)(卷积核3×3,padding=1,维度不变)
  • 第二次pool输出:(1,32,10,211)(池化核2×2,步长2,维度减半)
  • 扁平化后维度:32*10*211=67520,与代码自动计算结果一致,说明维度计算逻辑正确

3. 确保验证集DataLoader输入匹配训练集

检查valid_dl的输入形状是否与训练集一致,可打印一个batch的形状确认:

batch = next(iter(valid_dl))
print(batch[0].shape)  # 预期输出:(batch_size,1,40,844)

如果形状不符,可指定batch size为验证集样本数,避免维度变化:

valid_dl = learn.dls.test_dl(X_val, y_val, bs=21)

4. 重新初始化模型并训练

修正实例化错误后,重新执行模型初始化、训练流程,确保训练与推理使用同一模型结构:

model = DraftCNN()
learn = Learner(dls, model, loss_func=CrossEntropyLossFlat(), metrics=[accuracy, Precision(average='macro'), Recall(average='macro'), F1Score(average='macro')])
learn.fit_one_cycle(8)

内容的提问来源于stack exchange,提问作者Carlos Vega

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 03:37:04