如何快速加载大型NPZ文件的部分数据或随机样本?
Great question! When working with massive .npz files like your 180000×5×10×10 array, loading the entire dataset just to grab a small subset is a huge waste of time and memory. Here are two straightforward solutions to fix this:
1. Use Memory Mapping to Load Only Needed Slices
The key issue here is that numpy.load() by default loads the entire array into memory before you can slice it. Instead, use the mmap_mode parameter to create a memory-mapped array—this lets you access slices directly from disk without loading the full dataset first.
Here's how to implement it:
import numpy as np # Open the .npz file in read-only memory-mapped mode with np.load('data.npz', mmap_mode='r') as data: # Slice the array directly—only the required portion is read from disk subset = data['arr_0'][1:100, :, :, :]
mmap_mode='r'keeps the file read-only (safer for your original data) and maps the array to disk storage. When you slice it, numpy only fetches the specific blocks of data you need, which is way faster for large arrays.
2. Random Sampling Without Full Load
If you need random samples instead of a contiguous slice, memory mapping still works perfectly. Generate your random indices first, then pull only those samples from the memory-mapped array:
import numpy as np import random with np.load('data.npz', mmap_mode='r') as data: arr = data['arr_0'] total_samples = arr.shape[0] # Generate 100 random indices (adjust the number to your needs) random_indices = random.sample(range(total_samples), 100) # Extract the random subset—again, only these samples are loaded random_subset = arr[random_indices, :, :, :]
Quick Notes
- For most read-only tasks,
mmap_mode='r'is the best choice. Other options like'r+'(read-write) or'c'(copy-on-write) are available if you need to modify the array, but stick to'r'unless necessary. - SSD storage will make this even faster, but even with an HDD, this method is drastically quicker than loading the entire array.
- Keep the .npz file open (using the
withstatement or retaining thedataobject) if you need to pull multiple subsets—reopening the file repeatedly adds unnecessary overhead.
内容的提问来源于stack exchange,提问作者Lara

