PyTorch DecoderRNN报错:输入需为3维张量,实际得到2维
Hey there! Let's walk through why you're hitting that RuntimeError and fix up your code properly. First, the core error you see—input must have 3 dimensions, got 2—is just the tip of the iceberg; there are a few key issues with how you've set up the GRU and data flow.
Let's Break Down the Problems:
GRU Input Dimension Mismatch
PyTorch's GRU expects a 3D input tensor: either(seq_len, batch_size, input_size)or(batch_size, seq_len, input_size)(if you setbatch_first=True). Yourfeaturesare 2D ([10, 200]), missing the sequence length dimension.Incorrect GRU Initialization
You wrotenn.GRU(embed_size, hidden_size, hidden_size)—the third parameter here isnum_layers(number of GRU layers), nothidden_size. That's a mix-up that would cause all sorts of issues even if the dimension error was fixed.No Embedding for Captions
Yourcaptionsare raw word indices ([10,12]), which can't be fed directly into a GRU. You need an embedding layer to convert these indices into dense vector representations first.Wrong Use of Captions as Hidden State
You're passingcaptionsas the second argument to GRU, which is supposed to be the initial hidden state (hx). That's not what captions are for—they should be the sequence input to the GRU (after embedding), not the starting hidden state.Missing Linear Layer for Vocab Mapping
You referenceself.outin your forward pass, but never define it in__init__. This layer is needed to map the GRU's hidden outputs to your vocabulary size for prediction.
Here's the Fixed Code with Explanations:
import torch import torch.nn as nn class DecoderRNN(nn.Module): def __init__(self, embed_size, hidden_size, vocab_size, num_layers=1): super(DecoderRNN, self).__init__() self.hidden_size = hidden_size # Convert word indices to dense embeddings self.embedding = nn.Embedding(vocab_size, embed_size) # GRU with batch_first=True (easier to work with batch dimensions first) self.gru = nn.GRU(embed_size, hidden_size, num_layers=num_layers, batch_first=True) # Map GRU outputs to vocabulary size self.out = nn.Linear(hidden_size, vocab_size) # LogSoftmax for loss calculation (works with NLLLoss) self.softmax = nn.LogSoftmax(dim=2) def forward(self, features, captions): # Add a sequence dimension to image features: [10,200] → [10,1,200] features = features.unsqueeze(1) # Embed captions, and exclude the last token (usually <end> token) # Shape: [10,12] → [10,11,200] embeddings = self.embedding(captions[:, :-1]) # Concatenate image feature (as first sequence step) with caption embeddings # Shape becomes [10,12,200] inputs = torch.cat((features, embeddings), dim=1) # Pass through GRU: no need to pass initial hidden state (PyTorch uses zeros by default) output, hidden = self.gru(inputs) # Map to vocab and apply softmax output = self.softmax(self.out(output)) return output, hidden
Quick Shape Check:
- Input
features:[10,200]→ expanded to[10,1,200] - Input
captions:[10,12]→ embedded to[10,11,200](after slicing off last token) - Combined
inputs:[10,12,200](matches GRU's expected 3D shape) - Final
output:[10,12,vocab_size]→ perfect for comparing withcaptions[:,1:](your target sequence, missing the first<start>token) when calculating loss.
This should resolve the dimension error and all the other underlying issues in your original code.
内容的提问来源于stack exchange,提问作者Edamame

