You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何快速加载大型NPZ文件的部分数据或随机样本?

Efficiently Loading Subsets from Large .npz Files

Great question! When working with massive .npz files like your 180000×5×10×10 array, loading the entire dataset just to grab a small subset is a huge waste of time and memory. Here are two straightforward solutions to fix this:

1. Use Memory Mapping to Load Only Needed Slices

The key issue here is that numpy.load() by default loads the entire array into memory before you can slice it. Instead, use the mmap_mode parameter to create a memory-mapped array—this lets you access slices directly from disk without loading the full dataset first.

Here's how to implement it:

import numpy as np

# Open the .npz file in read-only memory-mapped mode
with np.load('data.npz', mmap_mode='r') as data:
    # Slice the array directly—only the required portion is read from disk
    subset = data['arr_0'][1:100, :, :, :]
  • mmap_mode='r' keeps the file read-only (safer for your original data) and maps the array to disk storage. When you slice it, numpy only fetches the specific blocks of data you need, which is way faster for large arrays.

2. Random Sampling Without Full Load

If you need random samples instead of a contiguous slice, memory mapping still works perfectly. Generate your random indices first, then pull only those samples from the memory-mapped array:

import numpy as np
import random

with np.load('data.npz', mmap_mode='r') as data:
    arr = data['arr_0']
    total_samples = arr.shape[0]
    # Generate 100 random indices (adjust the number to your needs)
    random_indices = random.sample(range(total_samples), 100)
    # Extract the random subset—again, only these samples are loaded
    random_subset = arr[random_indices, :, :, :]

Quick Notes

  • For most read-only tasks, mmap_mode='r' is the best choice. Other options like 'r+' (read-write) or 'c' (copy-on-write) are available if you need to modify the array, but stick to 'r' unless necessary.
  • SSD storage will make this even faster, but even with an HDD, this method is drastically quicker than loading the entire array.
  • Keep the .npz file open (using the with statement or retaining the data object) if you need to pull multiple subsets—reopening the file repeatedly adds unnecessary overhead.

内容的提问来源于stack exchange,提问作者Lara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:19:28