You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Conv2d输入维度错误求助及AlexNet集成模型结构修改需求

错误原因与解决方案

错误分析

RuntimeError: Expected 3D (unbatched) or 4D (batched) input to conv2d, but got input of size: [16, 18432] 是因为你添加的Conv2d层需要4D张量输入(格式为[批量大小, 通道数, 高度, 宽度]),但当前输入是2D张量,说明两个AlexNet的输出特征被错误展平,或拼接后丢失了空间维度。

同时按照你的需求,以下是移除两个AlexNet的avgpool和classifier模块,并修复维度问题的完整方案:

修正后的模型实现

import torch
import torch.nn as nn
from torchvision.models import alexnet

class MyEnsemble(nn.Module):
    def __init__(self, modelA, modelB):
        super().__init__()
        # 移除AlexNet的avgpool和classifier,仅保留特征提取部分
        self.modelA = nn.Sequential(*list(modelA.children())[:-2])
        self.modelB = nn.Sequential(*list(modelB.children())[:-2])
        
        # 拼接两个模型的输出通道(256+256=512),定义后续卷积层
        self.conv_layer = nn.Conv2d(512, 2, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))
        
        # 计算全连接层输入维度:假设输入图像为224x224,经过AlexNet特征提取后尺寸为13x13
        # 若你的输入图像尺寸不同,需重新计算该值(卷积后尺寸不变:13x13)
        self.classifier = nn.Linear(2 * 13 * 13, 2)

    def forward(self, x):
        # 提取两个模型的特征(保持4D张量格式)
        featA = self.modelA(x)  # 形状: [批量大小, 256, 13, 13]
        featB = self.modelB(x)  # 形状: [批量大小, 256, 13, 13]
        
        # 在通道维度拼接特征,得到[批量大小, 512, 13, 13]
        combined_feat = torch.cat([featA, featB], dim=1)
        
        # 卷积层处理拼接特征
        conv_out = self.conv_layer(combined_feat)  # 形状: [批量大小, 2, 13, 13]
        
        # 展平后输入全连接层
        flat_feat = torch.flatten(conv_out, 1)
        return self.classifier(flat_feat)

# 初始化预训练模型并创建集成模型
modelA = alexnet(pretrained=True)
modelB = alexnet(pretrained=True)
ensemble_model = MyEnsemble(modelA, modelB)

# 测试输入(批量16,3通道224x224图像)
test_input = torch.randn(16, 3, 224, 224)
output = ensemble_model(test_input)
print(output.shape)  # 预期输出: torch.Size([16, 2])

关键说明

  1. 移除冗余模块:通过nn.Sequential(*list(model.children())[:-2])截取AlexNet的features部分,直接丢弃avgpool和空的classifier模块。
  2. 保持维度正确:两个AlexNet输出的特征均为4D张量,拼接后仍为4D,满足Conv2d层的输入要求,避免了2D张量的错误输入。
  3. 调整全连接层维度:根据卷积层输出的空间尺寸计算全连接层的输入特征数,若你的输入图像尺寸不是224x224,需重新计算该值(可通过打印featA.shape查看具体的H和W)。

内容的提问来源于stack exchange,提问作者Ajay krishan gairola

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 00:32:53