You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch矩阵乘法形状不匹配RuntimeError问题排查求助

多输出线性回归模型的矩阵乘法形状匹配问题

问题背景

我是PyTorch新手,正在开发一个多输出线性回归模型,用于给单词上色(帮助 grapheme-color 联觉患者更轻松阅读)。模型输入为单词的编码向量:每个单词被表示为45维[0,1]浮点向量,其中(0,1]代表对应位置存在字母,0代表该位置无字母;输出为RGB值,每个样本对应[r值, g值, b值]。

报错信息

运行训练循环时出现如下错误:

RuntimeError: mat1 and mat2 shapes cannot be multiplied (90x1 and 45x3)

我推测需要对数据进行重塑,但不清楚具体操作方式与位置,尤其无法理解90x1矩阵的来源。

我的模型代码

class ColorPredictor(torch.nn.Module):
    #Constructor
    def __init__(self):
        super(ColorPredictor, self).__init__()
        self.linear = torch.nn.Linear(45, 3, device= device) #length of encoded word vectors & size of r,g,b vectors
        
    # Prediction
    def forward(self, x: torch.Tensor) -> torch.Tensor:
        y_pred = self.linear(x)
        return y_pred

数据加载代码

Dataset类

# Dataset Class
class Data(Dataset):
    # Constructor
    def __init__(self, inputs, outputs):
        self.x = inputs # a list of encoded word vectors
        self.y = outputs # a Pandas dataframe of r,g,b values converted to a torch tensor
        self.len = len(inputs)
    
    # Getter
    def __getitem__(self, index):
        return self.x[index], self.y[index]
    
    # Get number of samples
    def __len__(self):
        return self.len

训练/测试集拆分与DataLoader创建

# create train/test split
train_size = int(0.8 * len(data))
train_data = Data(inputs[:train_size], outputs[:train_size])
test_data = Data(inputs[train_size:], outputs[train_size:])
# create DataLoaders for training and testing sets
train_loader = DataLoader(dataset = train_data, batch_size=2)
test_loader = DataLoader(dataset = test_data, batch_size=2)

报错发生的训练循环

for epoch in range(epochs):
    # Train
    model.train() #training mode
    for x,y in train_loader:
        y_pred = model(x) #ERROR HERE
        loss = criterion(y_pred, y)
        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

后续尝试

我将45x1的输入张量改为2x45的输入张量(第二列全为0),首次遍历train_loader时可正常运行,但第二次遍历时再次出现矩阵乘法错误,报错为:

RuntimeError: mat1 and mat2 shapes cannot be multiplied (90x2 and 45x3)


解决方案

问题根源分析

  • 你的线性层torch.nn.Linear(45,3)要求输入张量的最后一维为45(对应每个样本的特征维度),但实际输入的张量形状不符合要求。
  • 90x1的来源:batch_size设为2,每个样本是45x1的二维向量,DataLoader拼接后得到形状为(2,45,1)的张量。PyTorch的Linear层会自动将除最后一维外的所有维度展平,于是(2,45,1)被展平成(90,1),和Linear层的权重矩阵(45,3)无法进行矩阵乘法(矩阵乘法要求第一个矩阵的列数等于第二个矩阵的行数)。

具体修复步骤

  1. 修正输入张量的形状
    修改Dataset类的__getitem__方法,确保返回的输入张量是一维(45,) 而非二维(45,1):

    def __getitem__(self, index):
        # 去除多余的维度,把(45,1)转为(45,)
        x_sample = torch.squeeze(self.x[index]) if self.x[index].dim() == 2 else self.x[index]
        y_sample = self.y[index]
        # 可选:确保y是一维(3,)张量
        y_sample = torch.squeeze(y_sample) if y_sample.dim() == 2 else y_sample
        return x_sample, y_sample
    

    或者在创建inputs列表时,提前对每个向量做处理:

    inputs = [torch.squeeze(vec) for vec in original_inputs]
    
  2. 验证输入形状
    在训练循环中添加打印语句,确认输入张量的形状是否符合预期(应为(2,45),对应batch_size=2,每个样本45维):

    for x,y in train_loader:
        print("Input shape:", x.shape)  # 预期输出: torch.Size([2, 45])
        y_pred = model(x)
        # 其余训练代码
    
  3. 纠正后续尝试的错误操作
    你之前将输入改为2x45是错误的——这相当于把一个样本的特征维度变成了2,完全不符合模型的设计。正确的输入形状应该是(batch_size, 45),每个样本对应45维特征。


内容的提问来源于stack exchange,提问作者A T

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 15:10:01