You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch中计算Embedding向量值梯度?双编码器LSTM训练报错解决

Fixing the "Embedding doesn't compute the gradient w.r.t. the indices" Error in PyTorch Dual Encoder LSTM

Hey there, let's break down what's causing this error and fix it step by step:

Root Cause

The AssertionError is triggered because you've set requires_grad=True on your input index tensors (context_matrix and response_matrix) in the training loop. The nn.Embedding layer is designed to only compute gradients for its own weight parameters, not for the input indices (which are just integer lookups). When you enable gradients on the indices, PyTorch hits that assertion check in the Embedding implementation.

Step-by-Step Fixes

1. Correct the Input Tensor Gradient Setting

First, remove the requires_grad=True flag from your input variables. By default, Variable (or torch.Tensor in newer PyTorch versions) has requires_grad=False for inputs, which is exactly what we need here:

# Training code modification
for i in range(epochs):
    context_matrix, response_matrix, y = make_matrices('/data/train_shuffled_onethousand.csv')
    # Remove requires_grad=True from input tensors
    context_matrix = autograd.Variable(context_matrix).cuda()
    response_matrix = autograd.Variable(response_matrix).cuda()
    y_label = autograd.Variable(y).cuda()  # Also wrap y in Variable for BCELoss compatibility
    
    y_preds = dual_encoder(context_matrix, response_matrix)
    loss = loss_func(y_preds, y_label)
    
    if i % 10 == 0:
        print("Epoch: ", i, ", Loss: ", loss.data[0])
    
    dual_encoder.zero_grad()
    loss.backward()
    torch.nn.utils.clip_grad_norm(dual_encoder.parameters(), 10)
    optimizer.step()

2. Fix Gradient Loss in DualEncoder's Forward Pass

Your current DualEncoder.forward method extracts .data from dot_var, which breaks the computation graph and prevents gradients from flowing back to your model parameters. This means even if you fixed the first error, your model wouldn't actually learn anything. Replace that loop with vectorized operations (which are faster and preserve gradients):

class DualEncoder(nn.Module):
    def __init__(self, encoder):
        super(DualEncoder, self).__init__()
        self.encoder = encoder
        self.number_of_layers = 1
        M = torch.FloatTensor(self.encoder.hidden_size, self.encoder.hidden_size).cuda()
        init.normal(M)
        self.M = nn.Parameter(M, requires_grad = True)
    
    def forward(self, contexts, responses):
        context_out, context_hn = self.encoder(contexts)
        response_out, response_hn = self.encoder(responses)
        
        # Get the final time step output from both encoders
        # Assuming input shape is (seq_len, batch_size) for LSTM with batch_first=False
        context_h = context_out[-1]  # Shape: (batch_size, hidden_size)
        response_h = response_out[-1]  # Shape: (batch_size, hidden_size)
        
        # Batch-wise score calculation using matrix multiplication
        # Expand M to match batch size, then compute (context_h @ M @ response_h.T) for each batch item
        scores = torch.bmm(
            context_h.unsqueeze(1), 
            torch.bmm(self.M.expand(context_h.size(0), -1, -1), response_h.unsqueeze(2))
        ).squeeze()
        
        y_preds = torch.sigmoid(scores)
        return y_preds

Your Encoder.__init__ has a few minor issues that could cause unexpected behavior:

  • self.dropout = 0, has an extra comma, turning it into a tuple instead of an integer. Remove the comma: self.dropout = 0
  • You're hardcoding self.num_layers = 1, self.num_directions = 1, and self.bidirectional = False instead of using the passed arguments. Use the parameters to keep consistency:
    def __init__(self, input_size, hidden_size, vocab_size, num_layers = 1, num_directions = 1, dropout = 0, bidirectional = False, rnn_type = 'lstm'):
        super(Encoder, self).__init__()
        self.input_size = input_size
        self.hidden_size = hidden_size
        self.vocab_size = vocab_size
        self.num_layers = num_layers  # Use passed argument
        self.num_directions = num_directions if bidirectional else 1
        self.dropout = dropout
        self.bidirectional = bidirectional
        # ... rest of your code
    

Why This Works

  • By setting requires_grad=False on the input indices, we respect the nn.Embedding layer's design, which only computes gradients for its weight matrix.
  • The vectorized operations in DualEncoder keep the computation graph intact, allowing gradients to flow back through the LSTM and Embedding layers during backpropagation.
  • Cleaning up the Encoder initialization ensures your model uses the parameters you intend, avoiding subtle bugs.

内容的提问来源于stack exchange,提问作者Janina Nuber

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:27:03