PyTorch人脸表情识别模型训练异常:替换测试加载器为训练加载器后精度不一致问题排查求助
Hey there, let's break down the problems causing the unexpected accuracy discrepancy when using trainloader in both training and evaluation phases, along with the extreme class accuracy results you're seeing.
Key Issues in Your Code
1. Incorrect Class Accuracy Calculation Logic
This is the primary reason behind the bizarre class accuracy outputs (all 0% except neutral at 100%). Let's look at your code snippet for counting correct predictions per class:
for i in range(len(results)): if results[i] == labels[i]: num_class[results[i]] = num_class[results[i]] + 1
You're incrementing the count for the predicted class when a match occurs, but you should be incrementing the count for the true class instead. The correct logic tracks how many samples of each actual class were correctly identified:
for i in range(len(results)): if results[i] == labels[i]: num_class[labels[i]] = num_class[labels[i]] + 1
Your current code only counts correct predictions under the predicted class label. Since your model is outputting almost exclusively "neutral" in evaluation mode, only true neutral samples contribute to a correct count—hence the 100% accuracy for neutral and 0% for all others.
2. Disabled Shuffling in Training DataLoader
Your training DataLoader has shuffle=False, which is a bad practice for training. Without shuffling, the model can easily memorize the order of training samples instead of learning generalizable features for facial expressions. Fix this by enabling shuffling:
trainloader = torch.utils.data.DataLoader(trainset, shuffle=True, batch_size=args.bn)
This helps the model learn more robust features and reduces issues with BatchNorm statistics (see next point).
3. Missing torch.no_grad() in Evaluation Phase
While you set model.eval() to disable training-specific layers like Dropout and BatchNorm, you didn't disable gradient computation during evaluation. This wastes memory and is unnecessary. Wrap your evaluation loop in a torch.no_grad() context:
model.eval() correct = 0 total = 0 total_loss = 0.0 sum_class = [0. for i in range(args.classnum)] num_class = [0. for i in range(args.classnum)] with torch.no_grad(): # Add this context manager for inputs, labels in trainloader: # Your existing evaluation code here
4. BatchNorm Behavior Differences Between Train/Eval Modes
When model.train() is active, BatchNorm updates its running mean and variance with each batch. In model.eval() mode, it uses the accumulated statistics from training. If your training data wasn't shuffled, batches might have had unbalanced class distributions, leading to inaccurate BatchNorm statistics. This can cause a noticeable drop in accuracy when switching to evaluation mode, which explains part of the discrepancy you saw between training and evaluation accuracy on the same dataset.
Additional Notes
- Your overall training accuracy (18%) is quite low, suggesting your model might need more training epochs, a better optimizer setup, or a more suitable architecture for facial expression recognition.
- Double-check your dataset balance: if one class (like neutral) is overrepresented, the model might bias towards predicting that class to minimize loss, which aligns with your evaluation results.
内容的提问来源于stack exchange,提问作者sir chen

