视频字幕生成任务中合并3D CNN与InceptionV2特征向量导致BLEU评分下降的原因咨询
It's such a frustrating common pitfall in multi-modal tasks: combining two seemingly complementary features ends up hurting performance instead of giving you a boost. Let’s break down why your current approach might be underperforming and walk through actionable fixes.
Key Reasons for the BLEU Score Drop
1. Your Projection Layers Are Untrained & Applied Offline
Looking at your code, you’ve defined projection layers for both feature types—but you’re using them with random initialization, no training whatsoever to process features offline. These layers haven’t learned to map raw 3D/2D features to a space that’s useful for captioning; they’re just applying arbitrary, unoptimized transformations that could distort or discard critical signal from your original features.
2. Simple Concatenation Isn’t Learning to Prioritize Features
Concatenation merges feature dimensions, but it doesn’t let your model weigh or prioritize what’s useful from each source:
- 3D CNN features capture spatiotemporal context (how objects move over frames)
- InceptionV2 features focus on fine-grained visual details (object shapes, colors)
Without a way to balance these signals, your captioning model might struggle to disentangle redundant or conflicting information from the merged vector.
3. BatchNorm Misuse in Offline Processing
Your projection layers include nn.BatchNorm1d, but when processing individual videos offline, each "batch" is just 15 frames. BatchNorm relies on global dataset statistics (mean/variance) to normalize features—using it on tiny per-video batches creates inconsistent feature distributions that don’t match what the layer expects, introducing unnecessary noise.
4. Over-Regularization Is Suppressing Useful Signal
The Dropout(0.5) in your projection layers applies heavy regularization before features even reach the captioning model. If your caption generator already has its own regularization, this double dose can smother the meaningful patterns in your features.
Fixes to Try (Ordered by Impact)
1. Move Projection & Fusion Into Your Captioning Model (End-to-End Training)
Instead of pre-processing and saving merged features, integrate these steps directly into your captioning pipeline so they can train alongside the decoder. This lets the projection layers learn to transform features specifically for your captioning task.
Here’s how to adjust your approach:
class VideoCaptionModel(nn.Module): def __init__(self, feature3d_dim=2048, feature2d_dim=1536, hidden_dim=512, vocab_size=...): super().__init__() # Projection layers (now part of the trainable captioning model) self.feature3d_proj = nn.Sequential( nn.Linear(feature3d_dim, 2 * hidden_dim), nn.ReLU(True), nn.Linear(2 * hidden_dim, hidden_dim) # Swap BatchNorm for LayerNorm (better for sequence data) nn.LayerNorm(hidden_dim) ) self.feature2d_proj = nn.Sequential( nn.Linear(feature2d_dim, 2 * hidden_dim), nn.ReLU(True), nn.Linear(2 * hidden_dim, hidden_dim), nn.LayerNorm(hidden_dim) ) # Gated fusion layer (learns to balance 3D/2D features) self.fusion_gate = nn.Linear(hidden_dim * 2, hidden_dim) # Rest of your captioning model (LSTM/Transformer decoder, etc.) self.decoder = ... def forward(self, f_3d, f_2d): # f_3d shape: (batch_size, 15, 2048) # f_2d shape: (batch_size, 15, 1536) proj_3d = self.feature3d_proj(f_3d) # (batch_size, 15, 512) proj_2d = self.feature2d_proj(f_2d) # (batch_size, 15, 512) # Gated fusion: dynamically weigh 3D vs 2D features concat_features = torch.cat([proj_3d, proj_2d], dim=-1) gate_weights = torch.sigmoid(self.fusion_gate(concat_features)) fused_features = gate_weights * proj_3d + (1 - gate_weights) * proj_2d # Pass fused features to your decoder output = self.decoder(fused_features) return output
2. Replace Concatenation with Smarter Fusion Strategies
If you want to stick with offline merging (though end-to-end is better), try these fusion methods instead of simple concatenation:
- Gated Fusion: Use a sigmoid gate to learn how much each feature contributes (like the example above)
- Element-Wise Addition/Multiplication: If projected to the same dimension, add or multiply features to combine signals without expanding the feature space
- Attention-Based Fusion: Let your decoder attend to both 3D and 2D features separately instead of merging upfront—this lets it pick relevant spatiotemporal/visual details for each word.
3. Fix Normalization or Remove It
- Swap
nn.BatchNorm1dfornn.LayerNorm: It normalizes across the feature dimension instead of batch, making it safe for offline sequence processing - If you keep BatchNorm, pre-train the projection layers on your full feature dataset first (using a proxy task like video category prediction) to compute correct running mean/variance, then freeze them for offline use.
4. Reduce Regularization
Remove the Dropout layer from your projection steps, or lower the rate to 0.2-0.3. Heavy dropout on already compressed features can destroy too much useful information.
5. Sanity Check: Test Raw Feature Concatenation
As a quick test, skip projection entirely and merge raw features:
# Merge raw features without projection combine_vectors = np.concatenate([f_3d, f_2d], axis=-1) # Shape (15, 3584)
If this performs better than your projected merged features, the problem is definitely with your untrained projection layers. If it still underperforms, your captioning model might need more capacity to handle the larger feature vector.
内容的提问来源于stack exchange,提问作者A_B_Y

