基于PyTorch自定义CNN:添加文本父类别信息优化多标签分类
Hey there! Your approach to adding parent category one-hot vectors as an extra input makes perfect sense for helping the model contextualize text sources. Let's break down the code changes step by step—PyTorch will absolutely handle backpropagation seamlessly once we wire everything up correctly.
First, let's fix a small bug in your original forward method (you had a missing loop over linear_layers), then add the parent input integration:
Step 1: Update the __init__ Method
We need to add a linear layer to process the parent category one-hot vectors. Let's add a new parameter parent_dim (the length of your one-hot vector) to the constructor, then define a linear layer to project this input into a feature space that plays nicely with your CNN's output.
Step 2: Modify the forward Method
The forward pass will now accept two inputs: the text tensor x and the parent one-hot tensor parent_input. We'll process the parent input, concatenate it with the CNN's extracted features, then feed the combined tensor through the rest of the network.
Here's the full modified code with comments explaining each change:
import torch import torch.nn as nn import torch.nn.functional as F class CNN(nn.Module): """Convolutional Neural Model used for training the models. The total number of kernels that will be used in this CNN is Co * len(Ks). Args: vocab_size: integer, number of words in the vocabulary emb_dim: integer, dimensionality of the word embeddings Co (number of filters): integer, stands for channels out and it is the number of kernels of the same size that will be used. Hu: list, list of integers specifying number of hidden units in each hidden layer. C: integer, number of units in the last layer (number of classes) Ks: list, list of integers specifying the size of the kernels to be used. parent_dim: integer, length of the parent category one-hot vector name: string, optional name suffix for the model """ def __init__(self, vocab_size, emb_dim, Co, Hu, C, Ks, parent_dim, name = 'generic'): super(CNN, self).__init__() self.num_embeddings = vocab_size self.embeddings_dim = emb_dim self.padding_index = 0 self.cnn_name = 'cnn_' + str(emb_dim) + '_' + str(Co) + '_' + str(Hu) + '_' + str(C) + '_' + str(Ks) + '_' + name self.Co = Co self.Hu = Hu self.C = C self.Ks = Ks # Original text processing layers self.embedding = nn.Embedding(self.num_embeddings, self.embeddings_dim, self.padding_index) self.convolutions = nn.ModuleList([nn.Conv2d(1,self.Co,(k, self.embeddings_dim)) for k in self.Ks]) self.relu = nn.ReLU() self.drop_out = nn.Dropout(p=0.5) # New: Linear layer to process parent category one-hot input # We'll project the one-hot vector to a dimension matching the CNN's feature output (Co * len(Ks)) # You can adjust this projection dimension if needed self.parent_proj = nn.Linear(parent_dim, Co * len(Ks)) # Adjust the first linear layer to accept concatenated features # Original input size was Co * len(Ks), now it's doubled (CNN features + projected parent features) cnn_feature_dim = Co * len(Ks) combined_input_dim = cnn_feature_dim * 2 units = [combined_input_dim] + Hu self.linear_layers = nn.ModuleList([nn.Linear(units[k],units[k+1]) for k in range(len(units)-1)]) self.linear_last = nn.Linear(self.Hu[-1], self.C) self.sigmoid = nn.Sigmoid() def forward(self, x, parent_input): # Process text input through CNN x = self.embedding(x) # Shape: (batch_size, seq_len, emb_dim) x = x.unsqueeze(1) # Add channel dimension for Conv2d: (batch_size, 1, seq_len, emb_dim) # Apply convolutions and pooling x = [self.relu(conv(x)).squeeze(3) for conv in self.convolutions] # Each element: (batch_size, Co, seq_len - k + 1) x = [F.max_pool1d(i, i.size(2)).squeeze(2) for i in x] # Each element: (batch_size, Co) cnn_features = torch.cat(x, 1) # Shape: (batch_size, Co * len(Ks)) # Process parent category one-hot input parent_features = self.relu(self.parent_proj(parent_input)) # Shape: (batch_size, Co * len(Ks)) # Concatenate CNN features and parent features combined_features = torch.cat([cnn_features, parent_features], dim=1) # Shape: (batch_size, 2 * Co * len(Ks)) # Pass through hidden linear layers x = combined_features for linear_layer in self.linear_layers: x = self.relu(linear_layer(x)) x = self.drop_out(x) # Final classification layer x = self.linear_last(x) x = self.sigmoid(x) return x
Key Notes to Keep in Mind
- Input Preparation: When training, you'll need to pass both your text tensor (shape
(batch_size, seq_len)) and parent one-hot tensor (shape(batch_size, parent_dim)) to the model. For example:# Sample input batch_text = torch.randint(0, vocab_size, (32, 50)) # 32 samples, 50-token sequences batch_parent = torch.randint(0, 2, (32, parent_dim)).float() # 32 one-hot vectors # Forward pass outputs = model(batch_text, batch_parent) - Projection Dimension: I chose to project the parent one-hot vector to match the CNN's feature dimension (
Co * len(Ks)) so their concatenation is balanced. You can tweak this dimension (e.g., use a smaller size) if you want the parent input to have less influence, or remove the projection entirely and concatenate the raw one-hot vector if its dimension is small. - Bug Fix: The original
forwardmethod had a missing loop overlinear_layers—I fixed that to ensure all hidden layers are used.
This setup lets the model learn to combine text features with parent category context, and PyTorch will automatically compute gradients for all layers (including the new parent_proj layer) during backpropagation.
内容的提问来源于stack exchange,提问作者Troy

