CNN新手:如何为ImageDataGenerator生成的样本分配标签?
Hey there! Since you're just starting out with CNNs and working on a leaf classification task, let's break down how to assign labels with ImageDataGenerator—super straightforward once you know the two main approaches that fit your use case.
1. Using flow_from_directory (Simplest for Your Current Setup)
This method is perfect because you already have separate folders for each leaf type. Here's how to set it up:
Step 1: Organize Your Folder Structure
First, restructure your image folders into a parent directory (let's call it leaf_dataset) with subfolders named after your classes. Like this:
leaf_dataset/ ├── leaf_A/ # Contains all images from image set X │ ├── img1.jpg │ ├── img2.jpg │ └── ... └── leaf_B/ # Contains all images from image set Y ├── img1.jpg ├── img2.jpg └── ...
The subfolder names (leaf_A and leaf_B) will automatically become your class labels.
Step 2: Code Implementation
Import the necessary modules and set up your ImageDataGenerator, then use flow_from_directory to load images and assign labels:
from tensorflow.keras.preprocessing.image import ImageDataGenerator # Initialize ImageDataGenerator (you can add augmentations here too!) datagen = ImageDataGenerator(rescale=1./255) # Normalize pixel values to 0-1 # Load training data train_generator = datagen.flow_from_directory( 'leaf_dataset/', # Path to your parent directory target_size=(224, 224), # Resize images to this size (adjust as needed) batch_size=32, class_mode='binary' # Use 'binary' for 2-class classification )
Step 3: Check Label Assignments
To confirm which class maps to which numeric label, you can print the class_indices attribute of the generator:
print(train_generator.class_indices) # Output will look like: {'leaf_A': 0, 'leaf_B': 1} (order depends on folder names)
This tells you that all images in leaf_A get labeled 0, and leaf_B get labeled 1 (or vice versa, based on alphabetical order of folder names).
2. Using flow_from_dataframe (More Flexible)
If you need more control over label assignment (e.g., custom numeric labels, or images aren't in class-named folders), this method works better.
Step 1: Create a DataFrame
First, make a pandas DataFrame that has two columns: one with the full path to each image, and another with its corresponding label. For example:
import pandas as pd import os # List all images and their labels image_paths = [] labels = [] # Add leaf_A images for img in os.listdir('path/to/leaf_A_folder'): image_paths.append(os.path.join('path/to/leaf_A_folder', img)) labels.append(0) # Assign custom label (e.g., 0 for leaf_A) # Add leaf_B images for img in os.listdir('path/to/leaf_B_folder'): image_paths.append(os.path.join('path/to/leaf_B_folder', img)) labels.append(1) # Assign custom label (e.g., 1 for leaf_B) # Create DataFrame df = pd.DataFrame({'image_path': image_paths, 'label': labels})
Step 2: Code Implementation
Use flow_from_dataframe to load images from the DataFrame:
train_generator = datagen.flow_from_dataframe( dataframe=df, x_col='image_path', # Column with image paths y_col='label', # Column with labels target_size=(224, 224), batch_size=32, class_mode='binary' )
This lets you explicitly set labels for each image, which is great if you need non-default label values or have a more complex file structure.
Quick Tips
- For binary classification (your case), use
class_mode='binary'. If you ever expand to more classes, switch toclass_mode='categorical'(for one-hot encoded labels) orclass_mode='sparse'(for integer labels). - Don't forget to normalize your pixel values (like
rescale=1./255) to help the model train better. - You can add data augmentations (e.g.,
rotation_range=20,width_shift_range=0.2) to the ImageDataGenerator to improve model generalization.
内容的提问来源于stack exchange,提问作者Ryan94

