You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效读取含3万张图片的文件夹并转换为20×20×3的NumPy数组?

高效处理大规模图片并转换为固定尺寸NumPy数组的方案

Hey, let's tackle this problem of processing 30k images efficiently without blowing up your memory or taking forever. Your current approach with a for-loop and Sci-Kit Image is hitting two main pain points: dynamic array resizing (which causes constant memory copying and bloat) and slower image loading compared to specialized libraries. Let's fix this step by step.

Why Your Current Approach Is Slow/Memory-Heavy

  • Every time you append to a NumPy array, it has to allocate a new, larger block of memory and copy all existing data over. Do this 30k times, and you're wasting tons of cycles on memory operations, not to mention the temporary memory spikes.
  • Sci-Kit Image is great for many tasks, but its image loading speed can't compete with libraries like OpenCV or Pillow, which are optimized for this exact use case.

Optimized Solutions

Since you know exactly how many images you have (~30k) and the target shape (20×20×3), we can preallocate a single array upfront. This eliminates all the dynamic resizing overhead and keeps memory usage completely predictable. For 30k images of uint8 pixels (0-255), this array will only take ~34MB of memory—nothing for modern systems.

Here's the code:

import os
import cv2
import numpy as np

# Configuration
image_dir = "/path/to/your/images"
target_size = (20, 20)
total_images = 30000  # Adjust if you count actual files first
num_channels = 3

# Preallocate array with uint8 (saves memory vs float32; convert later if needed)
dataset = np.zeros((total_images, target_size[0], target_size[1], num_channels), dtype=np.uint8)

# Get all valid image paths
image_paths = [
    os.path.join(image_dir, fname) 
    for fname in os.listdir(image_dir) 
    if fname.lower().endswith(('.png', '.jpg', '.jpeg'))
]

# Double-check image count matches (optional but safe)
assert len(image_paths) == total_images, f"Expected {total_images} images, found {len(image_paths)}"

# Process each image
for idx, img_path in enumerate(image_paths):
    # Load image (OpenCV uses BGR by default)
    img = cv2.imread(img_path)
    if img is None:
        print(f"Warning: Could not read {img_path}")
        continue
    
    # Resize to target size
    img_resized = cv2.resize(img, target_size)
    # Convert to RGB if needed (skip if BGR is fine for your use case)
    img_rgb = cv2.cvtColor(img_resized, cv2.COLOR_BGR2RGB)
    
    # Store in preallocated array (no memory copying here!)
    dataset[idx] = img_rgb

# Save the final array for later use
np.save("resized_images_20x20.npy", dataset)

2. Batch Processing (For Extreme Memory Constraints)

If for some reason even 34MB is too much (unlikely, but possible on super low-memory systems), you can process images in batches. Each batch is processed, saved to disk, then merged at the end:

import os
import cv2
import numpy as np

image_dir = "/path/to/your/images"
target_size = (20, 20)
batch_size = 1000  # Adjust based on your memory
image_paths = [
    os.path.join(image_dir, fname) 
    for fname in os.listdir(image_dir) 
    if fname.lower().endswith(('.png', '.jpg', '.jpeg'))
]
total_batches = len(image_paths) // batch_size + 1

# Process batches
for batch_idx in range(total_batches):
    start = batch_idx * batch_size
    end = min((batch_idx + 1) * batch_size, len(image_paths))
    batch_paths = image_paths[start:end]
    
    # Preallocate batch array
    batch_data = np.zeros((len(batch_paths), *target_size, 3), dtype=np.uint8)
    
    for idx, img_path in enumerate(batch_paths):
        img = cv2.imread(img_path)
        if img is None:
            print(f"Warning: Could not read {img_path}")
            continue
        img_resized = cv2.resize(img, target_size)
        batch_data[idx] = cv2.cvtColor(img_resized, cv2.COLOR_BGR2RGB)
    
    # Save batch to disk
    np.save(f"temp_batch_{batch_idx}.npy", batch_data)

# Merge all batches
all_batches = []
for batch_idx in range(total_batches):
    all_batches.append(np.load(f"temp_batch_{batch_idx}.npy"))
dataset = np.concatenate(all_batches, axis=0)
np.save("resized_images_20x20.npy", dataset)

# Clean up temporary files (optional)
for batch_idx in range(total_batches):
    os.remove(f"temp_batch_{batch_idx}.npy")

3. Multiprocessing (Speed Up IO-Bound Loading)

Image loading is IO-bound, so using multiple processes can cut down on total time. Here's how to do it with Python's multiprocessing module:

import os
import cv2
import numpy as np
from multiprocessing import Pool

def process_single_image(img_path):
    img = cv2.imread(img_path)
    if img is None:
        return None
    img_resized = cv2.resize(img, (20, 20))
    return cv2.cvtColor(img_resized, cv2.COLOR_BGR2RGB)

image_dir = "/path/to/your/images"
image_paths = [
    os.path.join(image_dir, fname) 
    for fname in os.listdir(image_dir) 
    if fname.lower().endswith(('.png', '.jpg', '.jpeg'))
]

# Use all CPU cores for processing
with Pool(processes=os.cpu_count()) as pool:
    results = pool.map(process_single_image, image_paths)

# Filter out any failed images
valid_results = [res for res in results if res is not None]
dataset = np.array(valid_results)
np.save("resized_images_20x20.npy", dataset)

Extra Tips

  • Stick to uint8: Unless you need floating-point values for processing, keep your array in uint8 format—it uses 4x less memory than float32.
  • Memory-Mapped Arrays: For truly massive datasets (100k+ images), use np.memmap to create an array that lives on disk instead of RAM. You can write to it incrementally without loading everything into memory.
  • Uniform Image Formats: If possible, convert all images to the same format (e.g., JPG) before processing—this speeds up loading since the library doesn't have to handle multiple decoders.

内容的提问来源于stack exchange,提问作者Dhruv Kapu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 16:42:43