You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Vision Transformer回归任务报错:矩阵形状不匹配如何解决?

问题分析

错误提示mat1 and mat2 shapes cannot be multiplied (32x1000 and 768x32)的核心原因:

  • 预训练的ViT-B/16默认输出是1000维的分类logits(对应ImageNet的1000类),而非我们需要的768维特征向量
  • 回归层错误地将输出维度设为num_classes * batch_size,这完全没必要,回归头只需输出对应任务的回归值数量(这里是1)
解决方法

以下两种方案都能解决问题,任选其一即可:

方案1:替换ViT原分类头为回归层

直接把ViT自带的1000类分类头替换成回归层,forward时直接输出回归结果:

import torch
import torch.nn as nn
import torch.optim as optim
from torchvision.models import vit_b_16

class RegressionViT(nn.Module):
    def __init__(self, num_classes=1, pretrained=True):
        super(RegressionViT, self).__init__()
        self.vit_b_16 = vit_b_16(pretrained=pretrained)
        # 替换原分类头为回归层:输入维度768,输出维度num_classes(这里是1)
        self.vit_b_16.heads = nn.Linear(self.vit_b_16.heads[0].in_features, num_classes)

    def forward(self, x):
        # 直接调用ViT的forward,此时输出就是回归值
        x = self.vit_b_16(x)
        return x

# 初始化模型
model = RegressionViT(num_classes=1)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)

criterion = nn.MSELoss()
optimizer = optim.Adam(model.parameters(), lr=0.0001)

方案2:提取ViT的特征向量后接回归层

保留ViT的encoder部分,在forward中手动提取cls token的特征,再传入回归层:

import torch
import torch.nn as nn
import torch.optim as optim
from torchvision.models import vit_b_16

class RegressionViT(nn.Module):
    def __init__(self, num_classes=1, pretrained=True):
        super(RegressionViT, self).__init__()
        self.vit_b_16 = vit_b_16(pretrained=pretrained)
        # 回归层:输入维度768(ViT cls token的特征维度),输出维度num_classes
        self.regressor = nn.Linear(self.vit_b_16.heads[0].in_features, num_classes)
        # 可选:冻结ViT encoder的参数,只训练回归层(如果不需要微调整个模型)
        # for param in self.vit_b_16.parameters():
        #     param.requires_grad = False

    def forward(self, x):
        # 提取ViT的中间特征:cls token的输出
        x = self.vit_b_16._process_input(x)
        n = x.shape[0]
        # 加上cls token和位置编码
        batch_class_token = self.vit_b_16.class_token.expand(n, -1, -1)
        x = torch.cat([batch_class_token, x], dim=1)
        x = x + self.vit_b_16.pos_embedding
        x = self.vit_b_16.dropout(x)
        # 经过encoder
        x = self.vit_b_16.encoder(x)
        # 取cls token的特征(第一个token)
        x = x[:, 0]
        # 传入回归层
        x = self.regressor(x)
        return x

# 初始化模型
model = RegressionViT(num_classes=1)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)

criterion = nn.MSELoss()
optimizer = optim.Adam(model.parameters(), lr=0.0001)
关键修改点说明
  • 移除了回归层中多余的batch_size乘数:回归头的输出维度只和任务需要的回归值数量有关,和batch size无关
  • 确保输入到回归层的是ViT的768维特征向量,而非预训练模型输出的1000维分类logits

内容的提问来源于stack exchange,提问作者SenthurLP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 16:57:38