You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch DecoderRNN报错:输入需为3维张量,实际得到2维

Fixing Your DecoderRNN for Image Captioning in PyTorch

Hey there! Let's walk through why you're hitting that RuntimeError and fix up your code properly. First, the core error you see—input must have 3 dimensions, got 2—is just the tip of the iceberg; there are a few key issues with how you've set up the GRU and data flow.

Let's Break Down the Problems:

  • GRU Input Dimension Mismatch
    PyTorch's GRU expects a 3D input tensor: either (seq_len, batch_size, input_size) or (batch_size, seq_len, input_size) (if you set batch_first=True). Your features are 2D ([10, 200]), missing the sequence length dimension.

  • Incorrect GRU Initialization
    You wrote nn.GRU(embed_size, hidden_size, hidden_size)—the third parameter here is num_layers (number of GRU layers), not hidden_size. That's a mix-up that would cause all sorts of issues even if the dimension error was fixed.

  • No Embedding for Captions
    Your captions are raw word indices ([10,12]), which can't be fed directly into a GRU. You need an embedding layer to convert these indices into dense vector representations first.

  • Wrong Use of Captions as Hidden State
    You're passing captions as the second argument to GRU, which is supposed to be the initial hidden state (hx). That's not what captions are for—they should be the sequence input to the GRU (after embedding), not the starting hidden state.

  • Missing Linear Layer for Vocab Mapping
    You reference self.out in your forward pass, but never define it in __init__. This layer is needed to map the GRU's hidden outputs to your vocabulary size for prediction.


Here's the Fixed Code with Explanations:

import torch
import torch.nn as nn

class DecoderRNN(nn.Module):
    def __init__(self, embed_size, hidden_size, vocab_size, num_layers=1):
        super(DecoderRNN, self).__init__()
        self.hidden_size = hidden_size
        
        # Convert word indices to dense embeddings
        self.embedding = nn.Embedding(vocab_size, embed_size)
        
        # GRU with batch_first=True (easier to work with batch dimensions first)
        self.gru = nn.GRU(embed_size, hidden_size, num_layers=num_layers, batch_first=True)
        
        # Map GRU outputs to vocabulary size
        self.out = nn.Linear(hidden_size, vocab_size)
        
        # LogSoftmax for loss calculation (works with NLLLoss)
        self.softmax = nn.LogSoftmax(dim=2)

    def forward(self, features, captions):
        # Add a sequence dimension to image features: [10,200] → [10,1,200]
        features = features.unsqueeze(1)
        
        # Embed captions, and exclude the last token (usually <end> token)
        # Shape: [10,12] → [10,11,200]
        embeddings = self.embedding(captions[:, :-1])
        
        # Concatenate image feature (as first sequence step) with caption embeddings
        # Shape becomes [10,12,200]
        inputs = torch.cat((features, embeddings), dim=1)
        
        # Pass through GRU: no need to pass initial hidden state (PyTorch uses zeros by default)
        output, hidden = self.gru(inputs)
        
        # Map to vocab and apply softmax
        output = self.softmax(self.out(output))
        
        return output, hidden

Quick Shape Check:

  • Input features: [10,200] → expanded to [10,1,200]
  • Input captions: [10,12] → embedded to [10,11,200] (after slicing off last token)
  • Combined inputs: [10,12,200] (matches GRU's expected 3D shape)
  • Final output: [10,12,vocab_size] → perfect for comparing with captions[:,1:] (your target sequence, missing the first <start> token) when calculating loss.

This should resolve the dimension error and all the other underlying issues in your original code.

内容的提问来源于stack exchange,提问作者Edamame

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:23:44