Keras训练提速:将分散图像合并为单训练/测试文件可行吗?
Absolutely! Merging your scattered images into a single file (like a numpy array) is a great way to cut down on training time with Keras, and it’s totally supported. Here’s a breakdown of why this works, how to do it, and how to train with the merged data:
The biggest bottleneck in your current training workflow is likely disk I/O, not GPU computation. Loading 60,000 individual small images means 60,000 separate hard disk read requests—waiting for these small, frequent operations adds up to significant wasted time. Merging into a single large file lets you load most or all of your data in just a few reads, eliminating the overhead of repeated small-file I/O and speeding up training noticeably.
- Numpy Array (Highly Recommended)
Images are numerical pixel data, and numpy arrays are binary-stored, compact, and lightning-fast to read/write. They also integrate seamlessly with Keras and TensorFlow. You can stack all training images into an array shaped(num_samples, height, width, channels)and store labels in a separate(num_samples,)array, then save them in a compressed format to save space. - CSV (Not Recommended for Images)
CSV is a text format, so storing image pixels would result in enormous file sizes (e.g., a 224x224 image has 50,176 values—60,000 of these would be 3 billion text characters). Reading and parsing this text is slow and inefficient, making CSV a poor choice for image data.
1. Preprocess and Merge Images into Numpy Files
First, iterate through your images, standardize them, and merge into numpy arrays:
import numpy as np from PIL import Image import os def merge_images_to_np(image_dir, label_list, img_size=(224, 224)): num_samples = len(label_list) # Initialize array for RGB images (adjust dtype/shape if using grayscale) images = np.empty((num_samples, img_size[0], img_size[1], 3), dtype=np.float32) for i, img_filename in enumerate(os.listdir(image_dir)): # Load and resize image img_path = os.path.join(image_dir, img_filename) img = Image.open(img_path).resize(img_size) # Normalize pixel values (adjust based on your preprocessing needs) images[i] = np.array(img) / 255.0 return images, np.array(label_list) # Assume your training images are in "train_images" folder, labels in train_labels list train_images, train_labels = merge_images_to_np("train_images", train_labels) # Save as compressed npz file (stores multiple arrays in one file) np.savez("train_data.npz", images=train_images, labels=train_labels) # Repeat for test set test_images, test_labels = merge_images_to_np("test_images", test_labels) np.savez("test_data.npz", images=test_images, labels=test_labels)
2. Train with Merged Data in Keras
Loading the merged files is straightforward, and training works almost identically to using individual images:
Scenario 1: Data fits entirely in memory
from keras.models import Sequential from keras.layers import Conv2D, MaxPooling2D, Flatten, Dense # Load training data train_data = np.load("train_data.npz") x_train = train_data["images"] y_train = train_data["labels"] # Load test data test_data = np.load("test_data.npz") x_test = test_data["images"] y_test = test_data["labels"] # Build your model (example CNN) model = Sequential([ Conv2D(32, (3,3), activation='relu', input_shape=(224,224,3)), MaxPooling2D((2,2)), Flatten(), Dense(128, activation='relu'), Dense(10, activation='softmax') ]) model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy']) # Train directly with the numpy arrays model.fit(x_train, y_train, epochs=10, batch_size=32, validation_data=(x_test, y_test))
Scenario 2: Data is too large for memory
Use tf.data.Dataset for batch-wise loading, or np.memmap for memory mapping:
import tensorflow as tf # Create a tf.data dataset from the numpy arrays (handles batch loading automatically) train_dataset = tf.data.Dataset.from_tensor_slices((x_train, y_train)) train_dataset = train_dataset.shuffle(1000).batch(32) test_dataset = tf.data.Dataset.from_tensor_slices((x_test, y_test)) test_dataset = test_dataset.batch(32) # Train with the dataset model.fit(train_dataset, epochs=10, validation_data=test_dataset)
If even loading the full numpy array is impossible, split your data into multiple smaller npz files and write a generator function to load batches from them during training.
- Preprocess Upfront: Do all preprocessing (resizing, normalization) before saving the numpy files—this avoids repeating work during training and saves more time. For online data augmentation, still use Keras'
ImageDataGeneratorortf.data'smapfunction. - Memory Calculation: For example, 60,000 224x224x3 float32 images take ~34GB of memory. If your system doesn't have enough, use batch loading or memory mapping.
- Backup Your Data: Merged files are single points of failure—make sure to back them up to avoid losing all your training data if the file gets corrupted.
内容的提问来源于stack exchange,提问作者Khanh Van Bui

