Vision Transformer回归任务报错:矩阵形状不匹配如何解决?
问题分析
错误提示mat1 and mat2 shapes cannot be multiplied (32x1000 and 768x32)的核心原因:
- 预训练的ViT-B/16默认输出是1000维的分类logits(对应ImageNet的1000类),而非我们需要的768维特征向量
- 回归层错误地将输出维度设为
num_classes * batch_size,这完全没必要,回归头只需输出对应任务的回归值数量(这里是1)
解决方法
以下两种方案都能解决问题,任选其一即可:
方案1:替换ViT原分类头为回归层
直接把ViT自带的1000类分类头替换成回归层,forward时直接输出回归结果:
import torch import torch.nn as nn import torch.optim as optim from torchvision.models import vit_b_16 class RegressionViT(nn.Module): def __init__(self, num_classes=1, pretrained=True): super(RegressionViT, self).__init__() self.vit_b_16 = vit_b_16(pretrained=pretrained) # 替换原分类头为回归层:输入维度768,输出维度num_classes(这里是1) self.vit_b_16.heads = nn.Linear(self.vit_b_16.heads[0].in_features, num_classes) def forward(self, x): # 直接调用ViT的forward,此时输出就是回归值 x = self.vit_b_16(x) return x # 初始化模型 model = RegressionViT(num_classes=1) device = torch.device("cuda" if torch.cuda.is_available() else "cpu") model.to(device) criterion = nn.MSELoss() optimizer = optim.Adam(model.parameters(), lr=0.0001)
方案2:提取ViT的特征向量后接回归层
保留ViT的encoder部分,在forward中手动提取cls token的特征,再传入回归层:
import torch import torch.nn as nn import torch.optim as optim from torchvision.models import vit_b_16 class RegressionViT(nn.Module): def __init__(self, num_classes=1, pretrained=True): super(RegressionViT, self).__init__() self.vit_b_16 = vit_b_16(pretrained=pretrained) # 回归层:输入维度768(ViT cls token的特征维度),输出维度num_classes self.regressor = nn.Linear(self.vit_b_16.heads[0].in_features, num_classes) # 可选:冻结ViT encoder的参数,只训练回归层(如果不需要微调整个模型) # for param in self.vit_b_16.parameters(): # param.requires_grad = False def forward(self, x): # 提取ViT的中间特征:cls token的输出 x = self.vit_b_16._process_input(x) n = x.shape[0] # 加上cls token和位置编码 batch_class_token = self.vit_b_16.class_token.expand(n, -1, -1) x = torch.cat([batch_class_token, x], dim=1) x = x + self.vit_b_16.pos_embedding x = self.vit_b_16.dropout(x) # 经过encoder x = self.vit_b_16.encoder(x) # 取cls token的特征(第一个token) x = x[:, 0] # 传入回归层 x = self.regressor(x) return x # 初始化模型 model = RegressionViT(num_classes=1) device = torch.device("cuda" if torch.cuda.is_available() else "cpu") model.to(device) criterion = nn.MSELoss() optimizer = optim.Adam(model.parameters(), lr=0.0001)
关键修改点说明
- 移除了回归层中多余的
batch_size乘数:回归头的输出维度只和任务需要的回归值数量有关,和batch size无关 - 确保输入到回归层的是ViT的768维特征向量,而非预训练模型输出的1000维分类logits
内容的提问来源于stack exchange,提问作者SenthurLP
相关产品推荐
相关产品推荐

