You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取ImageDataGenerator扩增图像的对应标签?兼议predict_generator优化

Solution: Get Bottleneck Features + Corresponding Labels in One Pass

Great question—this is a common pain point when doing data augmentation with bottleneck features, since you don’t want to run the generator twice (and end up with mismatched augmented samples and labels). Here’s how to solve it, both by modifying Keras’ source (if you’re comfortable) and with a safer custom function:

Option 1: Modify Keras’ training_generator.py (Direct Approach)

You’re exactly right about the core code change needed. Here’s the step-by-step breakdown:

  1. Locate the line where generator output is unpacked (around line 422 in the file you referenced):
    x, _ = generator_output
    
    Change it to preserve the labels:
    x, y = generator_output
    
  2. At the start of the predict_generator function, add a list to collect labels alongside features:
    all_labels = []
    
  3. Inside the loop where you append features to all_outs, add the batch labels to your new list:
    all_outs.append(to_list(x))
    all_labels.append(y)
    
  4. Finally, update the return statement to return both concatenated features and labels:
    return [np.concatenate(out) for out in all_outs], np.concatenate(all_labels)
    

Now when you call predict_generator, you’ll get both outputs in one pass:

bottleneck_features_train, train_labels = model.predict_generator(
    train_generator, 
    2 * nb_train_samples // batch_size
)

Option 2: Use a Custom Function (Safer, No Source Modification)

If you don’t want to tweak Keras’ core code, write a simple custom function to iterate over the generator and capture both features and labels simultaneously:

import numpy as np

def predict_with_labels(model, generator, steps):
    """Predict bottleneck features and capture their corresponding class labels."""
    all_features = []
    all_labels = []
    
    for _ in range(steps):
        # Fetch a batch of augmented images + their labels
        x_batch, y_batch = next(generator)
        # Generate bottleneck features for the batch
        features_batch = model.predict(x_batch, batch_size=x_batch.shape[0])
        # Store results
        all_features.append(features_batch)
        all_labels.append(y_batch)
    
    # Combine all batches into single arrays
    return np.concatenate(all_features), np.concatenate(all_labels)

# Example usage
bottleneck_features_train, train_labels = predict_with_labels(
    model, 
    train_generator, 
    2 * nb_train_samples // batch_size
)

Key Note About Consistency

Since your train_generator has shuffle="false", augmented samples will be generated in the order of your original directory—each original image will produce 2 augmented copies with the same label, and the labels array will align perfectly. If you ever use shuffle=True, set a fixed seed in flow_from_directory to ensure the augmentation sequence stays consistent (though shuffling is rarely needed during bottleneck feature extraction).

内容的提问来源于stack exchange,提问作者Dhiraj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:58:10