CNN新手求教:如何不使用ImageDataGenerator向Keras导入图像数据?
Hey there! As someone new to CNNs with Keras, I totally get wanting to understand how things work under the hood instead of just relying on ImageDataGenerator. Let's walk through exactly how to manually load your image datasets for that binary classification task, then cover how to compare performance with and without data augmentation.
First, Let's Confirm Your Folder Structure
I’m assuming your dataset is organized like this (super common for image classification tasks):
dataset/
├── train/
│ ├── class_1/
│ │ ├── img1.jpg
│ │ └── ...
│ └── class_2/
│ ├── img2.jpg
│ └── ...
└── val/
├── class_1/
└── class_2/
Step 1: Import Required Libraries
We’ll use standard tools for image loading, array handling, and label processing:
import os import numpy as np from PIL import Image from sklearn.preprocessing import LabelEncoder from tensorflow.keras.utils import to_categorical
Step 2: Write a Function to Load Images & Labels
This function will traverse your folder structure, convert images to numpy arrays, and process labels for your model:
def load_images_from_folder(folder, img_size=(224, 224)): images = [] labels = [] # Loop through each class folder for class_name in os.listdir(folder): class_path = os.path.join(folder, class_name) if not os.path.isdir(class_path): continue # Skip any non-folder items (like hidden files) # Load every image in the class folder for filename in os.listdir(class_path): img_path = os.path.join(class_path, filename) try: # Open image, convert to RGB (to avoid grayscale issues), resize img = Image.open(img_path).convert('RGB') img = img.resize(img_size) # Convert to numpy array and add to list img_array = np.array(img) images.append(img_array) labels.append(class_name) except Exception as e: print(f"Skipping corrupted image {img_path}: {str(e)}") # Convert lists to numpy arrays (required for Keras) images = np.array(images) labels = np.array(labels) # Encode labels from text (e.g., "class_1") to numerical values (0/1) label_encoder = LabelEncoder() encoded_labels = label_encoder.fit_transform(labels) # Convert to one-hot encoding (matches Keras' categorical class mode) categorical_labels = to_categorical(encoded_labels) return images, categorical_labels, label_encoder
Step 3: Load Training & Validation Data
Now use the function to load your datasets, and normalize pixel values (just like ImageDataGenerator’s rescale parameter):
# Match the image size you used with ImageDataGenerator IMG_SIZE = (224, 224) # Load training set train_imgs, train_labels, label_encoder = load_images_from_folder('dataset/train', img_size=IMG_SIZE) # Load validation set val_imgs, val_labels, _ = load_images_from_folder('dataset/val', img_size=IMG_SIZE) # Normalize pixel values to 0-1 range (critical for stable training) train_imgs = train_imgs / 255.0 val_imgs = val_imgs / 255.0
Step 4: Compare Performance With/Without Data Augmentation
Now you can train two identical models—one with your manually loaded (no augmentation) data, and one with ImageDataGenerator (with augmentation)—to compare accuracy.
First, Define a Shared CNN Model
We’ll use a simple, reusable model for fair comparison:
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense def build_simple_cnn(): model = Sequential([ Conv2D(32, (3,3), activation='relu', input_shape=(IMG_SIZE[0], IMG_SIZE[1], 3)), MaxPooling2D((2,2)), Conv2D(64, (3,3), activation='relu'), MaxPooling2D((2,2)), Flatten(), Dense(128, activation='relu'), Dense(2, activation='softmax') # Binary classification: 2 output classes ]) model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy']) return model
Train Model 1: No Data Augmentation (Manual Load)
model_no_aug = build_simple_cnn() history_no_aug = model_no_aug.fit( train_imgs, train_labels, epochs=10, batch_size=32, validation_data=(val_imgs, val_labels) )
Train Model 2: With Data Augmentation (ImageDataGenerator)
Use your existing generator setup, but make sure the target_size and rescale match your manual load:
from tensorflow.keras.preprocessing.image import ImageDataGenerator # Augmented training generator train_datagen = ImageDataGenerator( rescale=1./255, rotation_range=20, width_shift_range=0.2, height_shift_range=0.2, horizontal_flip=True ) # Validation generator (no augmentation, just normalization) val_datagen = ImageDataGenerator(rescale=1./255) train_generator = train_datagen.flow_from_directory( 'dataset/train', target_size=IMG_SIZE, batch_size=32, class_mode='categorical' ) val_generator = val_datagen.flow_from_directory( 'dataset/val', target_size=IMG_SIZE, batch_size=32, class_mode='categorical' ) # Train the model model_with_aug = build_simple_cnn() history_with_aug = model_with_aug.fit( train_generator, epochs=10, validation_data=val_generator )
Compare Results
You can pull the validation accuracy from both training histories to see the difference:
# Get final validation accuracy for no-augmentation model final_acc_no_aug = history_no_aug.history['val_accuracy'][-1] # Get final validation accuracy for augmented model final_acc_with_aug = history_with_aug.history['val_accuracy'][-1] print(f"Validation Accuracy (No Augmentation): {final_acc_no_aug:.4f}") print(f"Validation Accuracy (With Augmentation): {final_acc_with_aug:.4f}")
Quick Notes for Beginners
- Memory Considerations: If your dataset is huge, manual loading will use more RAM (since all images are loaded at once).
ImageDataGeneratorloads images in batches, which is gentler on memory. For large datasets, you could write a custom batch loader or usetf.data.Dataset(a bit more advanced, but worth learning later). - Label Flexibility: For binary classification, you can skip one-hot encoding and use
Dense(1, activation='sigmoid')withloss='binary_crossentropy'—this uses less memory and works just as well. - Consistency: Always match image size, normalization, and class mode between your manual load and generator setups to ensure a fair comparison.
内容的提问来源于stack exchange,提问作者Programmer

