如何获取ImageDataGenerator扩增图像的对应标签?兼议predict_generator优化
Great question—this is a common pain point when doing data augmentation with bottleneck features, since you don’t want to run the generator twice (and end up with mismatched augmented samples and labels). Here’s how to solve it, both by modifying Keras’ source (if you’re comfortable) and with a safer custom function:
Option 1: Modify Keras’ training_generator.py (Direct Approach)
You’re exactly right about the core code change needed. Here’s the step-by-step breakdown:
- Locate the line where generator output is unpacked (around line 422 in the file you referenced):
Change it to preserve the labels:x, _ = generator_outputx, y = generator_output - At the start of the
predict_generatorfunction, add a list to collect labels alongside features:all_labels = [] - Inside the loop where you append features to
all_outs, add the batch labels to your new list:all_outs.append(to_list(x)) all_labels.append(y) - Finally, update the return statement to return both concatenated features and labels:
return [np.concatenate(out) for out in all_outs], np.concatenate(all_labels)
Now when you call predict_generator, you’ll get both outputs in one pass:
bottleneck_features_train, train_labels = model.predict_generator( train_generator, 2 * nb_train_samples // batch_size )
Option 2: Use a Custom Function (Safer, No Source Modification)
If you don’t want to tweak Keras’ core code, write a simple custom function to iterate over the generator and capture both features and labels simultaneously:
import numpy as np def predict_with_labels(model, generator, steps): """Predict bottleneck features and capture their corresponding class labels.""" all_features = [] all_labels = [] for _ in range(steps): # Fetch a batch of augmented images + their labels x_batch, y_batch = next(generator) # Generate bottleneck features for the batch features_batch = model.predict(x_batch, batch_size=x_batch.shape[0]) # Store results all_features.append(features_batch) all_labels.append(y_batch) # Combine all batches into single arrays return np.concatenate(all_features), np.concatenate(all_labels) # Example usage bottleneck_features_train, train_labels = predict_with_labels( model, train_generator, 2 * nb_train_samples // batch_size )
Key Note About Consistency
Since your train_generator has shuffle="false", augmented samples will be generated in the order of your original directory—each original image will produce 2 augmented copies with the same label, and the labels array will align perfectly. If you ever use shuffle=True, set a fixed seed in flow_from_directory to ensure the augmentation sequence stays consistent (though shuffling is rarely needed during bottleneck feature extraction).
内容的提问来源于stack exchange,提问作者Dhiraj

