如何高效读取含3万张图片的文件夹并转换为20×20×3的NumPy数组?
Hey, let's tackle this problem of processing 30k images efficiently without blowing up your memory or taking forever. Your current approach with a for-loop and Sci-Kit Image is hitting two main pain points: dynamic array resizing (which causes constant memory copying and bloat) and slower image loading compared to specialized libraries. Let's fix this step by step.
Why Your Current Approach Is Slow/Memory-Heavy
- Every time you
appendto a NumPy array, it has to allocate a new, larger block of memory and copy all existing data over. Do this 30k times, and you're wasting tons of cycles on memory operations, not to mention the temporary memory spikes. - Sci-Kit Image is great for many tasks, but its image loading speed can't compete with libraries like OpenCV or Pillow, which are optimized for this exact use case.
Optimized Solutions
1. Preallocate a Fixed-Size NumPy Array (Most Recommended)
Since you know exactly how many images you have (~30k) and the target shape (20×20×3), we can preallocate a single array upfront. This eliminates all the dynamic resizing overhead and keeps memory usage completely predictable. For 30k images of uint8 pixels (0-255), this array will only take ~34MB of memory—nothing for modern systems.
Here's the code:
import os import cv2 import numpy as np # Configuration image_dir = "/path/to/your/images" target_size = (20, 20) total_images = 30000 # Adjust if you count actual files first num_channels = 3 # Preallocate array with uint8 (saves memory vs float32; convert later if needed) dataset = np.zeros((total_images, target_size[0], target_size[1], num_channels), dtype=np.uint8) # Get all valid image paths image_paths = [ os.path.join(image_dir, fname) for fname in os.listdir(image_dir) if fname.lower().endswith(('.png', '.jpg', '.jpeg')) ] # Double-check image count matches (optional but safe) assert len(image_paths) == total_images, f"Expected {total_images} images, found {len(image_paths)}" # Process each image for idx, img_path in enumerate(image_paths): # Load image (OpenCV uses BGR by default) img = cv2.imread(img_path) if img is None: print(f"Warning: Could not read {img_path}") continue # Resize to target size img_resized = cv2.resize(img, target_size) # Convert to RGB if needed (skip if BGR is fine for your use case) img_rgb = cv2.cvtColor(img_resized, cv2.COLOR_BGR2RGB) # Store in preallocated array (no memory copying here!) dataset[idx] = img_rgb # Save the final array for later use np.save("resized_images_20x20.npy", dataset)
2. Batch Processing (For Extreme Memory Constraints)
If for some reason even 34MB is too much (unlikely, but possible on super low-memory systems), you can process images in batches. Each batch is processed, saved to disk, then merged at the end:
import os import cv2 import numpy as np image_dir = "/path/to/your/images" target_size = (20, 20) batch_size = 1000 # Adjust based on your memory image_paths = [ os.path.join(image_dir, fname) for fname in os.listdir(image_dir) if fname.lower().endswith(('.png', '.jpg', '.jpeg')) ] total_batches = len(image_paths) // batch_size + 1 # Process batches for batch_idx in range(total_batches): start = batch_idx * batch_size end = min((batch_idx + 1) * batch_size, len(image_paths)) batch_paths = image_paths[start:end] # Preallocate batch array batch_data = np.zeros((len(batch_paths), *target_size, 3), dtype=np.uint8) for idx, img_path in enumerate(batch_paths): img = cv2.imread(img_path) if img is None: print(f"Warning: Could not read {img_path}") continue img_resized = cv2.resize(img, target_size) batch_data[idx] = cv2.cvtColor(img_resized, cv2.COLOR_BGR2RGB) # Save batch to disk np.save(f"temp_batch_{batch_idx}.npy", batch_data) # Merge all batches all_batches = [] for batch_idx in range(total_batches): all_batches.append(np.load(f"temp_batch_{batch_idx}.npy")) dataset = np.concatenate(all_batches, axis=0) np.save("resized_images_20x20.npy", dataset) # Clean up temporary files (optional) for batch_idx in range(total_batches): os.remove(f"temp_batch_{batch_idx}.npy")
3. Multiprocessing (Speed Up IO-Bound Loading)
Image loading is IO-bound, so using multiple processes can cut down on total time. Here's how to do it with Python's multiprocessing module:
import os import cv2 import numpy as np from multiprocessing import Pool def process_single_image(img_path): img = cv2.imread(img_path) if img is None: return None img_resized = cv2.resize(img, (20, 20)) return cv2.cvtColor(img_resized, cv2.COLOR_BGR2RGB) image_dir = "/path/to/your/images" image_paths = [ os.path.join(image_dir, fname) for fname in os.listdir(image_dir) if fname.lower().endswith(('.png', '.jpg', '.jpeg')) ] # Use all CPU cores for processing with Pool(processes=os.cpu_count()) as pool: results = pool.map(process_single_image, image_paths) # Filter out any failed images valid_results = [res for res in results if res is not None] dataset = np.array(valid_results) np.save("resized_images_20x20.npy", dataset)
Extra Tips
- Stick to uint8: Unless you need floating-point values for processing, keep your array in uint8 format—it uses 4x less memory than float32.
- Memory-Mapped Arrays: For truly massive datasets (100k+ images), use
np.memmapto create an array that lives on disk instead of RAM. You can write to it incrementally without loading everything into memory. - Uniform Image Formats: If possible, convert all images to the same format (e.g., JPG) before processing—this speeds up loading since the library doesn't have to handle multiple decoders.
内容的提问来源于stack exchange,提问作者Dhruv Kapu

