序列标注任务中词/字符n-gram特征提取的CNN架构选型咨询
Hey there! I totally get the confusion when stepping into CNNs for text tasks—especially when you’re used to other sequence models. Let’s unpack your questions one by one to make this clearer.
First: What does MaxPool1D actually do in this setup?
You’re right that Conv1D is perfect for grabbing n-gram features: a Conv1D layer with kernel size k will slide over your sequence and extract local patterns from every consecutive k characters/words, turning each into a feature map.
Now for MaxPool1D—it doesn’t just spit out a single max value unless you use global pooling (which is for classification, not sequence tagging). Here’s the breakdown based on how you configure it:
- If you use a
pool_sizeof, say, 2 withstrides=2and no padding: For each chunk of 2 consecutive positions in your feature map, it takes the maximum value across each feature dimension. So if your input after Conv1D is(batch_size, seq_len, num_filters), the output becomes(batch_size, seq_len//2, num_filters)—you’re halving the sequence length while keeping all feature dimensions. - If you set
padding='same'andstrides=1withpool_size=2: The output sequence length stays the same as the input. For each position, it looks at the current and next position, then takes the max for each feature dimension. This keeps your sequence aligned with the original input (critical for sequence tagging!) while smoothing out noise and highlighting the strongest local features. - Global MaxPool1D (pool_size equal to sequence length): This does output a single vector of
(batch_size, num_filters)—but this is only useful for tasks like text classification where you need a fixed-length representation of the whole sequence. Don’t use this for sequence tagging, since you need to predict a label for every input token.
What’s a reasonable architecture for sequence tagging?
Since you need to output a label per token, your priority is to preserve the sequence length throughout most of the model. Here’s a solid starting point:
- Input Layer: Combine character-level embeddings and word-level embeddings (if you have both) into a single input tensor of shape
(batch_size, seq_len, total_embed_dim). - Conv1D Layers: Stack 1-3 Conv1D layers with
padding='same'(so sequence length stays the same). Use multiple kernel sizes (e.g., 2, 3, 4) in separate branches, then concatenate their outputs—this lets you capture n-grams of different lengths at once. For example:# Example using Keras conv2 = Conv1D(filters=64, kernel_size=2, padding='same', activation='relu')(input_embeds) conv3 = Conv1D(filters=64, kernel_size=3, padding='same', activation='relu')(input_embeds) conv4 = Conv1D(filters=64, kernel_size=4, padding='same', activation='relu')(input_embeds) combined = concatenate([conv2, conv3, conv4], axis=-1) - Optional Pooling (If Needed): If you want to add some pooling without shrinking the sequence, use MaxPool1D with
padding='same'andstrides=1. Or skip pooling entirely—many sequence tagging CNNs just use stacked Conv1D layers with same padding. - Final Prediction Layers: Add a
Conv1D(filters=num_tags, kernel_size=1)layer (this acts like a dense layer applied to each token) to map the features to your tag space. Then wrap this in a CRF layer—this is a game-changer for sequence tagging, as it models the dependencies between consecutive tags (e.g., "B-PER" is almost always followed by "I-PER" instead of "B-LOC"). - Training: Use categorical cross-entropy (if not using CRF) or the CRF loss function, and optimize with Adam.
Quick Pro Tips
- Character vs Word Embeddings: Character-level CNNs shine at capturing morphological patterns (like prefixes/suffixes in English), while word-level embeddings bring semantic context. Combining both usually gives better results than using one alone.
- Filter Counts: Start with 64-128 filters per Conv1D layer—you can scale up if your dataset is large.
- Regularization: Add Dropout layers between Conv1D layers or after embeddings to prevent overfitting, especially if your dataset is small.
内容的提问来源于stack exchange,提问作者Igor Shalyminov

