You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PyTorch与多层感知器的情感分析任务中Embedding维度不匹配报错排查

Fixing "mat1 and mat2 shapes cannot be multiplied" Error in PyTorch Sentiment Analysis MLP

Hey there, let's break down why you're getting that shape mismatch error and fix your code step by step.

The error mat1 and mat2 shapes cannot be multiplied (50176x100 and 25002x256) tells us exactly what's wrong: your first linear layer expects inputs of size INPUT_DIM (25002) but is receiving embedded tensors of size EMBEDDING_DIM (100) after the embedding layer. On top of that, there are a few other logic issues in your forward pass and layer setup.

Key Issues in Your Code

  • Wrong input dimension for the first Linear layer: Your MLP starts with nn.Linear(self.INPUT_DIM, self.HIDDEN_DIM), but after embedding, each sample's feature size is EMBEDDING_DIM, not INPUT_DIM.
  • Confused forward pass: You're calling self.model(embedded) and then immediately self.model(x)—this is redundant and incorrect, since x is the raw token index tensor, not the embedded feature tensor.
  • 3D to 2D conversion missing: The nn.Embedding layer outputs a 3D tensor ([sequence_length, batch_size, embedding_dim] or [batch_size, sequence_length, embedding_dim] depending on your input layout), but MLPs expect 2D tensors ([batch_size, feature_dim]). You need to collapse the sequence dimension (e.g., take the mean over the sequence length to get a single embedding per sample).
  • Variable order issue: You're using INPUT_DIM, EMBEDDING_DIM etc. before assigning them as instance variables in the class constructor—this would throw an undefined variable error if those aren't global.
  • Incorrect activation function placement: Putting ReLU after the final Linear layer and before Sigmoid is unnecessary for binary classification; ReLU could zero out values that Sigmoid needs to produce a valid probability.

Fixed Code with Explanations

import torch.nn as nn

class MultilayerPerceptron(nn.Module):
    def __init__(self, input_dim, embedding_dim, hidden_dim, output_dim):
        # Initialize superclass first
        super(MultilayerPerceptron, self).__init__()
        
        # Define instance variables first to avoid undefined errors
        self.input_dim = input_dim
        self.embedding_dim = embedding_dim
        self.hidden_dim = hidden_dim
        self.output_dim = output_dim
        
        # Embedding layer: maps token indices to dense vectors
        self.embedding = nn.Embedding(self.input_dim, self.embedding_dim)
        
        # MLP layers: input is now embedding_dim (not input_dim)
        self.model = nn.Sequential(
            # First linear layer: takes embedded features (flattened or averaged)
            nn.Linear(self.embedding_dim, self.hidden_dim),
            nn.ReLU(),  # ReLU on hidden layer to introduce non-linearity
            nn.Linear(self.hidden_dim, self.output_dim),
            nn.Sigmoid()  # Sigmoid for binary classification (outputs 0-1 probability)
        )

    def forward(self, x):
        # x shape: [batch_size, sequence_length] (adjust if your layout is [seq_len, batch])
        embedded = self.embedding(x)  # shape: [batch_size, sequence_length, embedding_dim]
        
        # Collapse sequence dimension: take mean over sequence length to get one embedding per sample
        # If your input is [sequence_length, batch_size], use embedded.mean(dim=0) instead
        embedded_mean = embedded.mean(dim=1)  # shape: [batch_size, embedding_dim]
        
        # Pass the averaged embedding through the MLP
        output = self.model(embedded_mean)
        
        return output

# Initialize hyperparameters
INPUT_DIM = len(TEXT.vocab)
EMBEDDING_DIM = 100
HIDDEN_DIM = 256
OUTPUT_DIM = 1

# Create model instance with all required parameters
model = MultilayerPerceptron(INPUT_DIM, EMBEDDING_DIM, HIDDEN_DIM, OUTPUT_DIM)

What We Fixed

  1. Constructor parameters: We now pass all required dimensions (embedding_dim, output_dim) to the class constructor instead of relying on global variables, making the class reusable.
  2. Linear layer input fixed: The first Linear layer now takes embedding_dim as input, matching the output of the embedding layer after averaging.
  3. Forward pass cleaned up: We only process the embedded tensor, convert it to 2D using mean pooling over the sequence length, then pass it through the MLP.
  4. Activation order fixed: ReLU is applied after the hidden layer to add non-linearity, and Sigmoid is the final layer for binary classification probabilities.
  5. Variable order corrected: Instance variables are defined before using them in layers, preventing undefined variable errors.

内容的提问来源于stack exchange,提问作者Flavio Spadavecchia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 18:07:49