You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

向Numpy数组追加1.5万+图片后出现内存错误求助

Fixing Memory Error When Loading Large Image Datasets for CNN Training

Hey Vidit, let's break down how to solve this memory issue—sparse matrices actually aren't the right fit here, and I'll explain why plus give you practical, actionable alternatives.

Why Sparse Matrices Won't Help

Sparse matrices only save memory when most of your data is zero-valued. For standard RGB (or even grayscale) images, almost every pixel has a non-zero value (0-255 range), so storing them as sparse matrices would add overhead (tracking non-zero indices) instead of saving space. We need approaches tailored specifically to image datasets.

Practical Solutions to Avoid Memory Overload

1. Use Data Generators (Load On-Demand)

Instead of loading all 90k images into a numpy array upfront, load batches of images as you train your CNN. Most deep learning frameworks have built-in tools for this, which keeps memory usage low and integrates smoothly with training loops.

Keras Example with ImageDataGenerator

from tensorflow.keras.preprocessing.image import ImageDataGenerator

# Initialize generator with preprocessing (resize to 512x512)
datagen = ImageDataGenerator(rescale=1./255)

# Load images directly from a directory (organize data into class subfolders first)
train_generator = datagen.flow_from_directory(
    'path/to/your/image/directory',
    target_size=(512, 512),
    batch_size=32,  # Adjust based on your available memory
    class_mode='categorical'  # Use 'binary' for binary classification tasks
)

# Train your model using the generator
model.fit(train_generator, epochs=10, steps_per_epoch=len(train_generator))

PyTorch Example with Dataset and DataLoader

from torch.utils.data import Dataset, DataLoader
from PIL import Image
import os
from torchvision import transforms

class CustomImageDataset(Dataset):
    def __init__(self, img_dir, transform=None):
        self.img_dir = img_dir
        self.transform = transform
        self.img_paths = [os.path.join(img_dir, f) for f in os.listdir(img_dir) if f.endswith(('.png', '.jpg'))]

    def __len__(self):
        return len(self.img_paths)

    def __getitem__(self, idx):
        img_path = self.img_paths[idx]
        image = Image.open(img_path).convert('RGB')
        if self.transform:
            image = self.transform(image)
        # Add your label logic here if needed
        return image

# Define preprocessing pipeline (resize to 512x512)
transform = transforms.Compose([
    transforms.Resize((512, 512)),
    transforms.ToTensor()
])

# Create dataset and loader
dataset = CustomImageDataset('path/to/your/image/directory', transform=transform)
dataloader = DataLoader(dataset, batch_size=32, shuffle=True)

# Iterate over batches during training
for batch in dataloader:
    # Train on the current batch
    pass

2. Memory-Mapped Numpy Arrays

If you still want to preprocess all images upfront but avoid loading everything into RAM, use numpy.memmap. This stores the array on disk and lets you access only the parts you need during training.

import numpy as np
from PIL import Image
import os

# Calculate total size: 90000 images * 512x512 pixels * 3 channels (RGB) * 1 byte (uint8)
total_size = (90000, 512, 512, 3)
dtype = np.uint8  # Use uint8 instead of float to save massive memory

# Create a memory-mapped file
memmap_arr = np.memmap('images_memmap.npy', dtype=dtype, mode='w+', shape=total_size)

# Preprocess and write images to the memmap array
for i, img_path in enumerate(os.listdir('path/to/images')):
    img = Image.open(img_path).convert('RGB').resize((512, 512))
    memmap_arr[i] = np.array(img)

# Flush changes to disk
memmap_arr.flush()

# Later, load the memmap array for training (only loads needed parts into RAM)
loaded_memmap = np.memmap('images_memmap.npy', dtype=dtype, mode='r', shape=total_size)

# Access batches like this:
batch = loaded_memmap[0:32]  # Only this batch is loaded into memory

3. Split Preprocessed Data into Batches

Preprocess your images in chunks, save each chunk as a separate numpy file, then load them one at a time during training:

import numpy as np
from PIL import Image
import os

img_paths = [os.path.join('path/to/images', f) for f in os.listdir('path/to/images') if f.endswith(('.png', '.jpg'))]
batch_size = 1000  # Adjust based on your memory capacity

for i in range(0, len(img_paths), batch_size):
    batch_paths = img_paths[i:i+batch_size]
    batch_arr = []
    for path in batch_paths:
        img = Image.open(path).convert('RGB').resize((512, 512))
        batch_arr.append(np.array(img))
    # Save batch to disk
    np.save(f'preprocessed_batch_{i//batch_size}.npy', np.array(batch_arr))

# During training, load batches one by one
for batch_file in os.listdir('path/to/preprocessed_batches'):
    batch = np.load(os.path.join('path/to/preprocessed_batches', batch_file))
    # Train on this batch
    pass

4. Reduce Data Precision

Images are naturally stored as 8-bit integers (0-255). If you're converting them to float64 unnecessarily, switch back to uint8 (for preprocessing) or use float32 instead of float64 for model inputs. This cuts memory usage by 75% (float64 → uint8) or 50% (float64 → float32).

Final Notes

Data generators are almost always the best approach for large image datasets—they keep memory usage low and work seamlessly with training workflows. Memmap or batch-split methods are useful if you need preprocessed data, but generators are more efficient for most cases.

内容的提问来源于stack exchange,提问作者vidit02100

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:11:53