基于PyTorch与多层感知器的情感分析任务中Embedding维度不匹配报错排查
Fixing "mat1 and mat2 shapes cannot be multiplied" Error in PyTorch Sentiment Analysis MLP
Hey there, let's break down why you're getting that shape mismatch error and fix your code step by step.
The error mat1 and mat2 shapes cannot be multiplied (50176x100 and 25002x256) tells us exactly what's wrong: your first linear layer expects inputs of size INPUT_DIM (25002) but is receiving embedded tensors of size EMBEDDING_DIM (100) after the embedding layer. On top of that, there are a few other logic issues in your forward pass and layer setup.
Key Issues in Your Code
- Wrong input dimension for the first Linear layer: Your MLP starts with
nn.Linear(self.INPUT_DIM, self.HIDDEN_DIM), but after embedding, each sample's feature size isEMBEDDING_DIM, notINPUT_DIM. - Confused forward pass: You're calling
self.model(embedded)and then immediatelyself.model(x)—this is redundant and incorrect, sincexis the raw token index tensor, not the embedded feature tensor. - 3D to 2D conversion missing: The
nn.Embeddinglayer outputs a 3D tensor ([sequence_length, batch_size, embedding_dim]or[batch_size, sequence_length, embedding_dim]depending on your input layout), but MLPs expect 2D tensors ([batch_size, feature_dim]). You need to collapse the sequence dimension (e.g., take the mean over the sequence length to get a single embedding per sample). - Variable order issue: You're using
INPUT_DIM,EMBEDDING_DIMetc. before assigning them as instance variables in the class constructor—this would throw an undefined variable error if those aren't global. - Incorrect activation function placement: Putting ReLU after the final Linear layer and before Sigmoid is unnecessary for binary classification; ReLU could zero out values that Sigmoid needs to produce a valid probability.
Fixed Code with Explanations
import torch.nn as nn class MultilayerPerceptron(nn.Module): def __init__(self, input_dim, embedding_dim, hidden_dim, output_dim): # Initialize superclass first super(MultilayerPerceptron, self).__init__() # Define instance variables first to avoid undefined errors self.input_dim = input_dim self.embedding_dim = embedding_dim self.hidden_dim = hidden_dim self.output_dim = output_dim # Embedding layer: maps token indices to dense vectors self.embedding = nn.Embedding(self.input_dim, self.embedding_dim) # MLP layers: input is now embedding_dim (not input_dim) self.model = nn.Sequential( # First linear layer: takes embedded features (flattened or averaged) nn.Linear(self.embedding_dim, self.hidden_dim), nn.ReLU(), # ReLU on hidden layer to introduce non-linearity nn.Linear(self.hidden_dim, self.output_dim), nn.Sigmoid() # Sigmoid for binary classification (outputs 0-1 probability) ) def forward(self, x): # x shape: [batch_size, sequence_length] (adjust if your layout is [seq_len, batch]) embedded = self.embedding(x) # shape: [batch_size, sequence_length, embedding_dim] # Collapse sequence dimension: take mean over sequence length to get one embedding per sample # If your input is [sequence_length, batch_size], use embedded.mean(dim=0) instead embedded_mean = embedded.mean(dim=1) # shape: [batch_size, embedding_dim] # Pass the averaged embedding through the MLP output = self.model(embedded_mean) return output # Initialize hyperparameters INPUT_DIM = len(TEXT.vocab) EMBEDDING_DIM = 100 HIDDEN_DIM = 256 OUTPUT_DIM = 1 # Create model instance with all required parameters model = MultilayerPerceptron(INPUT_DIM, EMBEDDING_DIM, HIDDEN_DIM, OUTPUT_DIM)
What We Fixed
- Constructor parameters: We now pass all required dimensions (
embedding_dim,output_dim) to the class constructor instead of relying on global variables, making the class reusable. - Linear layer input fixed: The first Linear layer now takes
embedding_dimas input, matching the output of the embedding layer after averaging. - Forward pass cleaned up: We only process the embedded tensor, convert it to 2D using mean pooling over the sequence length, then pass it through the MLP.
- Activation order fixed: ReLU is applied after the hidden layer to add non-linearity, and Sigmoid is the final layer for binary classification probabilities.
- Variable order corrected: Instance variables are defined before using them in layers, preventing undefined variable errors.
内容的提问来源于stack exchange,提问作者Flavio Spadavecchia
相关产品推荐
相关产品推荐

