Keras中LSTM验证准确率与自行测算结果不符问题咨询
Hey there! Let's break down why you're seeing such a big gap between your model's reported validation accuracy (0.30) and the manual calculation using model.predict_generator (~0.18). This is almost always tied to how you're handling the time series data generators or matching predictions to labels—let's go through the most likely culprits:
1. Shuffled Validation Generator (The #1 Culprit)
If your validation TimeseriesGenerator has shuffle=True, that's a huge problem. When shuffle is on, the generator mixes up the order of time series sequences, which means the predictions from predict_generator will be in a random order, while your original validation labels are still in their original sequence. Comparing these mismatched lists will give you a garbage accuracy score.
Fix: Always set shuffle=False for your validation generator.
2. Mismatched Generator Parameters
Double-check that your training and validation generators use identical core parameters:
timesteps: Same number of time steps per sequencesampling_rate: Same interval between sampled time stepsstride: Same step size between consecutive sequencesbatch_size: Consistent batch size (though less critical for accuracy, it can affect edge cases)
If any of these differ, your model is being validated on a different sequence structure than what you're using for manual prediction.
3. Incorrect Label Matching with Time Series Sequences
Remember: TimeseriesGenerator creates sequences of timesteps length, and each sequence is paired with the label that comes immediately after the sequence. For example, if you have data points [t0, t1, t2, t3] with timesteps=2, the generator will produce:
- Sequence
[t0, t1]→ labelt2 - Sequence
[t1, t2]→ labelt3
When you calculate accuracy manually, you need to make sure you're comparing predictions to these "post-sequence" labels, not the full original label array. If you're using the entire y_val array instead of the subset that matches the generator's sequences, you'll be comparing the wrong labels.
Quick Check: Verify the number of samples from your validation generator matches the number of predictions:
val_gen = sequence.TimeseriesGenerator(X_val, y_val, timesteps, shuffle=False) predictions = model.predict_generator(val_gen) print(f"Number of predictions: {len(predictions)}") print(f"Expected number of validation sequences: {len(y_val) - timesteps + 1}")
These numbers should be identical.
4. Label Format Mismatch
- If your model uses
categorical_crossentropy(one-hot encoded labels), make sure you're converting predictions to class indices correctly withnp.argmax(predictions, axis=1), and that your true labels are also converted from one-hot to indices. - If you're using
sparse_categorical_crossentropy(integer labels), skip thenp.argmaxstep for true labels.
Mixing up these formats will lead to incorrect accuracy calculations.
5. Metric Misalignment
Double-check the accuracy metric you used when compiling your model. For example:
- If you compiled with
metrics=['categorical_accuracy']but your manual calculation uses sparse accuracy logic (comparing integer predictions to integer labels), you'll see a mismatch. - Ensure your manual calculation uses the same logic as the model's built-in metric.
Quick Test Code to Debug
Here's a snippet to help you align predictions and labels correctly:
# Initialize validation generator with shuffle=False val_gen = sequence.TimeseriesGenerator(X_val, y_val, timesteps, sampling_rate=1, stride=1, shuffle=False) # Get predictions and convert to class indices pred_probs = model.predict_generator(val_gen) pred_classes = np.argmax(pred_probs, axis=1) # Extract true labels from the generator (matches the sequence order) true_classes = [] for _, batch_labels in val_gen: # Adjust this line based on your label format (one-hot vs integer) true_classes.extend(np.argmax(batch_labels, axis=1)) # For one-hot labels # true_classes.extend(batch_labels) # For integer labels # Calculate manual accuracy manual_accuracy = np.mean(pred_classes == np.array(true_classes)) print(f"Manual Accuracy: {manual_accuracy:.4f}") print(f"Model Reported Validation Accuracy: {model.history.history['val_accuracy'][-1]:.4f}")
Run this code, and the two accuracy values should now be close (if not identical). If they're still off, double-check your train/validation split to ensure you didn't accidentally include training data in the validation set, or vice versa.
内容的提问来源于stack exchange,提问作者mr.gratitude

