如何在PyTorch中实现可学习嵌入矩阵以完成高维向量低维压缩?
Alright, let's walk through how to build this learnable embedding matrix E that maps your 4096-dimensional vector f to the 20-dimensional theta vector. This is straightforward with PyTorch's built-in modules, and I'll cover two common, practical approaches below.
Step 1: Clarify the Matrix Dimensions First
To avoid mix-ups, let's lock in the math:
- Your input
fis a 4096×1 vector - The embedding matrix
Eneeds to be 20×4096 (since matrix multiplicationE * foutputs a 20×1 vector, which is yourtheta) - PyTorch is optimized for batch inputs, so we'll adjust dimensions slightly to fit that workflow without breaking your original formula.
Step 2: Approach 1 - Use nn.Linear (Simplest & Recommended)
The nn.Linear layer is perfect here: its weight matrix is exactly our E, and it's automatically set as a learnable parameter by default. We'll disable the bias term since your formula doesn't include one.
import torch from torch import nn # Define your model class class EmbeddingReducer(nn.Module): def __init__(self, input_dim=4096, output_dim=20): super().__init__() # Linear layer with no bias (matches theta = E*f) self.embed = nn.Linear(input_dim, output_dim, bias=False) def forward(self, f): # Handle the (4096,1) input shape by squeezing the extra dimension if f.dim() == 2 and f.shape[1] == 1: f = f.squeeze(dim=1) # Converts to (4096,) theta = self.embed(f) # Reshape back to (20,1) to match your original formula's output shape theta = theta.unsqueeze(dim=1) return theta # Initialize the model model = EmbeddingReducer() # Test with your sample input f = torch.randn(4096, 1) theta = model(f) print(f"Theta shape: {theta.shape}") # Should output torch.Size([20, 1])
Step 3: Approach 2 - Explicitly Define nn.Parameter
If you want direct, explicit control over the matrix E, you can define it as a learnable nn.Parameter tensor instead:
class EmbeddingReducer(nn.Module): def __init__(self, input_dim=4096, output_dim=20): super().__init__() # Initialize E as a 20x4096 learnable parameter # We use randn for initial weights (standard practice for embeddings) self.E = nn.Parameter(torch.randn(output_dim, input_dim)) def forward(self, f): # Perform direct matrix multiplication E * f theta = torch.matmul(self.E, f) return theta # Initialize and test model = EmbeddingReducer() f = torch.randn(4096, 1) theta = model(f) print(f"Theta shape: {theta.shape}") # torch.Size([20, 1])
Step 4: Training Setup to Learn E
To make E update during training, follow standard PyTorch training steps—PyTorch will automatically track gradients for both approaches above:
# Choose an optimizer (Adam is a safe default for embedding learning) optimizer = torch.optim.Adam(model.parameters(), lr=1e-3) # Dummy training loop (replace with your actual dataset and task-specific loss) for epoch in range(10): optimizer.zero_grad() # Reset gradients each epoch # Get your input f (this could come from a DataLoader in real use) f = torch.randn(4096, 1) # Forward pass to get theta theta = model(f) # Dummy loss (replace this with your actual loss function) loss = theta.sum() # Backpropagate the loss to update E loss.backward() optimizer.step() print(f"Epoch {epoch+1}, Loss: {loss.item():.4f}")
Quick Notes:
- Both methods make
Elearnable—PyTorch handles gradient tracking fornn.Linearweights andnn.Parametertensors automatically. - If you're working with batches of
fvectors (e.g.,(batch_size, 4096)), just adjust your input shape and the forward pass will handle it seamlessly.
内容的提问来源于stack exchange,提问作者jason

