MLP手写数字分类异常求助:新增类别2无法识别但标签互换后功能反转
Alright, let's break down exactly what's going on here—since swapping labels proves your model can learn the new class (2) flawlessly, the problem isn't with the class itself, but with how the model interacts with the original class 0's data or your network setup.
Possible Root Causes
Data Distribution/Quality Disparity
The label swap result makes this the most likely culprit. Think through these questions:- Do you have far fewer samples of digit 0 compared to digit 2? Class imbalance can make the model prioritize learning the majority class and ignore the minority one.
- Are digit 0 samples noisier, blurry, or inconsistently preprocessed? Maybe your normalization step works for 2 but fails for 0 (e.g., some 0 samples have pixel values outside the intended [0,1] range).
- When you visualize samples, do 0s have way more variability (like thick vs. thin strokes, shifted positions) that the model can't generalize from?
Sigmoid Activation Saturation
Your MLP uses two layers of Sigmoid, which is infamous for gradient vanishing when inputs are too large or too small. Here's the breakdown:- If digit 0's features push most Sigmoid units into their saturated regions (close to 0 or 1), the gradients during backpropagation become nearly zero. This means the model can't update its weights to learn 0's unique patterns.
- Digit 2's features might fall right into Sigmoid's linear, gradient-friendly range (around the midpoint 0.5), so the model picks up those features quickly—even when relabeled as class 0.
Mismatched Loss Function/Task Setup
When you added digit 2, did you switch from binary classification (0 vs 1) to multi-class classification (0 vs 1 vs 2)? If you kept using a binary cross-entropy loss with a single Sigmoid output, that's a critical mismatch. Multi-class tasks require a Softmax output layer and categorical cross-entropy loss. Even for binary test runs (2 vs 0), double-check that your label encoding (e.g., one-hot vs. scalar) aligns with your loss function.
Fixes to Try
Audit Your Data First
- Count samples per class to fix imbalance (oversample 0, undersample 2, or use weighted loss to prioritize the underrepresented class).
- Visualize 10-20 samples of each digit to spot preprocessing errors or noisy data.
- Ensure all samples are normalized identically (e.g., divide pixel values by 255 to lock them in the [0,1] range).
Replace Sigmoid with More Robust Activations
Swap Sigmoid for ReLU or LeakyReLU in your hidden layers. These activations avoid saturation and keep gradients flowing, making it easier for the model to learn all classes. If you must keep Sigmoid, standardize your input features (subtract the mean, divide by standard deviation) to keep inputs in the [-5,5] range where Sigmoid is most responsive.Adjust Network Initialization
Use Xavier initialization for your weights instead of random small values. This ensures initial inputs to Sigmoid units stay in a range that avoids immediate saturation, giving the model a better starting point to learn all classes.Verify Loss Function & Output Layer
For multi-class classification (3 digits), set your output layer to use Softmax activation and pair it with categorical cross-entropy loss. For binary test runs, confirm your labels are encoded correctly (e.g., 0 and 1 for the two classes) and match the binary cross-entropy loss.
内容的提问来源于stack exchange,提问作者GGS

