训练全卷积网络时如何忽略指定标签?Caffe+Cityscapes场景实操问询
Hey there! Let's break this down into practical steps and underlying principles so you can get your FCN training properly with Cityscapes.
Practical Implementation in Caffe
The key to telling your network to ignore label 255 lies in configuring the loss layer correctly, plus a few checks to ensure your data pipeline is set up right.
1. Configure the SoftmaxWithLoss Layer
Caffe's SoftmaxWithLoss layer has a built-in parameter ignore_label that's exactly designed for this scenario. You just need to add this parameter to your prototxt file's loss layer definition:
layer { name: "segmentation_loss" type: "SoftmaxWithLoss" bottom: "fcn_score" # Your FCN's final output feature map bottom: "label" # Your ground truth label map (with 255 for ignored regions) top: "loss" loss_param { ignore_label: 255 normalize: true # Optional: Normalize loss by the number of valid (non-255) pixels } }
What this does: During training, any pixel with a label of 255 will be skipped entirely—its loss won't be calculated, and no gradients will be propagated back to the network for that pixel.
2. Verify Your Data Pipeline
Make sure your data layer (like Data or ImageData) reads the label maps correctly:
- Ensure labels are loaded as unsigned 8-bit integers (
UINT8) so the value 255 is preserved (don't normalize labels to 0-1 or scale them down). - Double-check that your category ID conversion script has correctly mapped all ignored Cityscapes classes (like
unlabeled,ego vehicle,rectification border) to 255.
3. Set the Correct Number of Output Channels
Your FCN's final convolution layer should output only the number of valid training classes (for Cityscapes, that's typically 19 foreground classes plus 1 background class, total 20). Remember: 255 is a separate "ignore" flag, not part of these target classes.
Underlying Principles
Let's dive into why this works and why it's necessary:
How Loss Calculation Works Normally
Semantic segmentation uses per-pixel cross-entropy loss. For each pixel, the network outputs a probability distribution over your target classes. The loss measures how far this distribution is from the one-hot encoded ground truth label (where only the correct class has a probability of 1).
What ignore_label Does
When you set ignore_label:255, Caffe's loss layer adds a check before calculating loss for each pixel:
- If the pixel's label is 255, it skips that pixel entirely. No loss is computed, and no gradients are generated for that pixel's predictions.
- This effectively tells the network: "These regions don't contain useful training information—don't waste effort learning to predict them."
Why We Need This for Cityscapes
Cityscapes includes regions that are either:
- Unlabeled: Areas where annotators couldn't reliably label the scene (like distant objects, blurry regions).
- Irrelevant: Regions like the ego vehicle's own body, which we don't want the model to predict.
Ignoring these prevents the model from learning incorrect patterns from noisy or irrelevant data, which would hurt its performance on the classes we care about.
Quick Check for Validation
When evaluating your model (calculating mIoU, for example), make sure you also ignore pixels labeled 255—only compute metrics on the valid, labeled regions. This gives you an accurate measure of how well the model performs on the classes that matter.
内容的提问来源于stack exchange,提问作者MeanStreet

