PyTorch中计算Embedding向量值梯度?双编码器LSTM训练报错解决
Hey there, let's break down what's causing this error and fix it step by step:
Root Cause
The AssertionError is triggered because you've set requires_grad=True on your input index tensors (context_matrix and response_matrix) in the training loop. The nn.Embedding layer is designed to only compute gradients for its own weight parameters, not for the input indices (which are just integer lookups). When you enable gradients on the indices, PyTorch hits that assertion check in the Embedding implementation.
Step-by-Step Fixes
1. Correct the Input Tensor Gradient Setting
First, remove the requires_grad=True flag from your input variables. By default, Variable (or torch.Tensor in newer PyTorch versions) has requires_grad=False for inputs, which is exactly what we need here:
# Training code modification for i in range(epochs): context_matrix, response_matrix, y = make_matrices('/data/train_shuffled_onethousand.csv') # Remove requires_grad=True from input tensors context_matrix = autograd.Variable(context_matrix).cuda() response_matrix = autograd.Variable(response_matrix).cuda() y_label = autograd.Variable(y).cuda() # Also wrap y in Variable for BCELoss compatibility y_preds = dual_encoder(context_matrix, response_matrix) loss = loss_func(y_preds, y_label) if i % 10 == 0: print("Epoch: ", i, ", Loss: ", loss.data[0]) dual_encoder.zero_grad() loss.backward() torch.nn.utils.clip_grad_norm(dual_encoder.parameters(), 10) optimizer.step()
2. Fix Gradient Loss in DualEncoder's Forward Pass
Your current DualEncoder.forward method extracts .data from dot_var, which breaks the computation graph and prevents gradients from flowing back to your model parameters. This means even if you fixed the first error, your model wouldn't actually learn anything. Replace that loop with vectorized operations (which are faster and preserve gradients):
class DualEncoder(nn.Module): def __init__(self, encoder): super(DualEncoder, self).__init__() self.encoder = encoder self.number_of_layers = 1 M = torch.FloatTensor(self.encoder.hidden_size, self.encoder.hidden_size).cuda() init.normal(M) self.M = nn.Parameter(M, requires_grad = True) def forward(self, contexts, responses): context_out, context_hn = self.encoder(contexts) response_out, response_hn = self.encoder(responses) # Get the final time step output from both encoders # Assuming input shape is (seq_len, batch_size) for LSTM with batch_first=False context_h = context_out[-1] # Shape: (batch_size, hidden_size) response_h = response_out[-1] # Shape: (batch_size, hidden_size) # Batch-wise score calculation using matrix multiplication # Expand M to match batch size, then compute (context_h @ M @ response_h.T) for each batch item scores = torch.bmm( context_h.unsqueeze(1), torch.bmm(self.M.expand(context_h.size(0), -1, -1), response_h.unsqueeze(2)) ).squeeze() y_preds = torch.sigmoid(scores) return y_preds
3. Clean Up Encoder Class Initialization (Optional but Recommended)
Your Encoder.__init__ has a few minor issues that could cause unexpected behavior:
self.dropout = 0,has an extra comma, turning it into a tuple instead of an integer. Remove the comma:self.dropout = 0- You're hardcoding
self.num_layers = 1,self.num_directions = 1, andself.bidirectional = Falseinstead of using the passed arguments. Use the parameters to keep consistency:def __init__(self, input_size, hidden_size, vocab_size, num_layers = 1, num_directions = 1, dropout = 0, bidirectional = False, rnn_type = 'lstm'): super(Encoder, self).__init__() self.input_size = input_size self.hidden_size = hidden_size self.vocab_size = vocab_size self.num_layers = num_layers # Use passed argument self.num_directions = num_directions if bidirectional else 1 self.dropout = dropout self.bidirectional = bidirectional # ... rest of your code
Why This Works
- By setting
requires_grad=Falseon the input indices, we respect thenn.Embeddinglayer's design, which only computes gradients for its weight matrix. - The vectorized operations in
DualEncoderkeep the computation graph intact, allowing gradients to flow back through the LSTM and Embedding layers during backpropagation. - Cleaning up the Encoder initialization ensures your model uses the parameters you intend, avoiding subtle bugs.
内容的提问来源于stack exchange,提问作者Janina Nuber

