You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CNN新手求教:如何不使用ImageDataGenerator向Keras导入图像数据?

Hey there! As someone new to CNNs with Keras, I totally get wanting to understand how things work under the hood instead of just relying on ImageDataGenerator. Let's walk through exactly how to manually load your image datasets for that binary classification task, then cover how to compare performance with and without data augmentation.

First, Let's Confirm Your Folder Structure

I’m assuming your dataset is organized like this (super common for image classification tasks):

dataset/
├── train/
│ ├── class_1/
│ │ ├── img1.jpg
│ │ └── ...
│ └── class_2/
│ ├── img2.jpg
│ └── ...
└── val/
├── class_1/
└── class_2/

Step 1: Import Required Libraries

We’ll use standard tools for image loading, array handling, and label processing:

import os
import numpy as np
from PIL import Image
from sklearn.preprocessing import LabelEncoder
from tensorflow.keras.utils import to_categorical

Step 2: Write a Function to Load Images & Labels

This function will traverse your folder structure, convert images to numpy arrays, and process labels for your model:

def load_images_from_folder(folder, img_size=(224, 224)):
    images = []
    labels = []
    
    # Loop through each class folder
    for class_name in os.listdir(folder):
        class_path = os.path.join(folder, class_name)
        if not os.path.isdir(class_path):
            continue  # Skip any non-folder items (like hidden files)
        
        # Load every image in the class folder
        for filename in os.listdir(class_path):
            img_path = os.path.join(class_path, filename)
            try:
                # Open image, convert to RGB (to avoid grayscale issues), resize
                img = Image.open(img_path).convert('RGB')
                img = img.resize(img_size)
                # Convert to numpy array and add to list
                img_array = np.array(img)
                images.append(img_array)
                labels.append(class_name)
            except Exception as e:
                print(f"Skipping corrupted image {img_path}: {str(e)}")
    
    # Convert lists to numpy arrays (required for Keras)
    images = np.array(images)
    labels = np.array(labels)
    
    # Encode labels from text (e.g., "class_1") to numerical values (0/1)
    label_encoder = LabelEncoder()
    encoded_labels = label_encoder.fit_transform(labels)
    # Convert to one-hot encoding (matches Keras' categorical class mode)
    categorical_labels = to_categorical(encoded_labels)
    
    return images, categorical_labels, label_encoder

Step 3: Load Training & Validation Data

Now use the function to load your datasets, and normalize pixel values (just like ImageDataGenerator’s rescale parameter):

# Match the image size you used with ImageDataGenerator
IMG_SIZE = (224, 224)

# Load training set
train_imgs, train_labels, label_encoder = load_images_from_folder('dataset/train', img_size=IMG_SIZE)
# Load validation set
val_imgs, val_labels, _ = load_images_from_folder('dataset/val', img_size=IMG_SIZE)

# Normalize pixel values to 0-1 range (critical for stable training)
train_imgs = train_imgs / 255.0
val_imgs = val_imgs / 255.0

Step 4: Compare Performance With/Without Data Augmentation

Now you can train two identical models—one with your manually loaded (no augmentation) data, and one with ImageDataGenerator (with augmentation)—to compare accuracy.

First, Define a Shared CNN Model

We’ll use a simple, reusable model for fair comparison:

from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense

def build_simple_cnn():
    model = Sequential([
        Conv2D(32, (3,3), activation='relu', input_shape=(IMG_SIZE[0], IMG_SIZE[1], 3)),
        MaxPooling2D((2,2)),
        Conv2D(64, (3,3), activation='relu'),
        MaxPooling2D((2,2)),
        Flatten(),
        Dense(128, activation='relu'),
        Dense(2, activation='softmax')  # Binary classification: 2 output classes
    ])
    model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])
    return model

Train Model 1: No Data Augmentation (Manual Load)

model_no_aug = build_simple_cnn()
history_no_aug = model_no_aug.fit(
    train_imgs, train_labels,
    epochs=10,
    batch_size=32,
    validation_data=(val_imgs, val_labels)
)

Train Model 2: With Data Augmentation (ImageDataGenerator)

Use your existing generator setup, but make sure the target_size and rescale match your manual load:

from tensorflow.keras.preprocessing.image import ImageDataGenerator

# Augmented training generator
train_datagen = ImageDataGenerator(
    rescale=1./255,
    rotation_range=20,
    width_shift_range=0.2,
    height_shift_range=0.2,
    horizontal_flip=True
)

# Validation generator (no augmentation, just normalization)
val_datagen = ImageDataGenerator(rescale=1./255)

train_generator = train_datagen.flow_from_directory(
    'dataset/train',
    target_size=IMG_SIZE,
    batch_size=32,
    class_mode='categorical'
)

val_generator = val_datagen.flow_from_directory(
    'dataset/val',
    target_size=IMG_SIZE,
    batch_size=32,
    class_mode='categorical'
)

# Train the model
model_with_aug = build_simple_cnn()
history_with_aug = model_with_aug.fit(
    train_generator,
    epochs=10,
    validation_data=val_generator
)

Compare Results

You can pull the validation accuracy from both training histories to see the difference:

# Get final validation accuracy for no-augmentation model
final_acc_no_aug = history_no_aug.history['val_accuracy'][-1]
# Get final validation accuracy for augmented model
final_acc_with_aug = history_with_aug.history['val_accuracy'][-1]

print(f"Validation Accuracy (No Augmentation): {final_acc_no_aug:.4f}")
print(f"Validation Accuracy (With Augmentation): {final_acc_with_aug:.4f}")

Quick Notes for Beginners

  • Memory Considerations: If your dataset is huge, manual loading will use more RAM (since all images are loaded at once). ImageDataGenerator loads images in batches, which is gentler on memory. For large datasets, you could write a custom batch loader or use tf.data.Dataset (a bit more advanced, but worth learning later).
  • Label Flexibility: For binary classification, you can skip one-hot encoding and use Dense(1, activation='sigmoid') with loss='binary_crossentropy'—this uses less memory and works just as well.
  • Consistency: Always match image size, normalization, and class mode between your manual load and generator setups to ensure a fair comparison.

内容的提问来源于stack exchange,提问作者Programmer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:23:05