使用numpy.load切片加载文件:仅加载指定部分还是全量加载后切片?
np.load('myfile.npy')[start:end] load only the sliced portion or the entire file? Great question — this is a super common point of confusion when working with large NumPy datasets, so let’s break it down clearly:
- Your current code loads the entire file into memory first, then slices it.
Here’s the breakdown: When you call np.load('myfile.npy') without any special flags, NumPy reads the full contents of the .npy file into a fully in-memory array. The [start:end] slice operation happens after the entire array is already loaded — it just creates a view (or copy, depending on the slice type) of the in-memory array, not a subset pulled directly from the file.
If you’re dealing with large files where loading everything would hog too much RAM, you need to use memory mapping to load only the slice you need. Here’s the proper approach:
# Load only the [start:end] slice directly from the file arr_slice = np.load('myfile.npy', mmap_mode='r')[start:end]
When you use mmap_mode='r', NumPy creates a memory-mapped array that references the file on disk instead of loading the whole thing into RAM. When you slice this mapped array, only the specific range [start:end] gets loaded into memory. This is way more memory-efficient for big datasets.
A quick few notes about memory mapping:
- The mapped array is read-only with
mmap_mode='r'(ideal for just reading slices). If you need to modify the array, you can usemmap_mode='r+'instead — just be careful not to accidentally overwrite your original file. - Memory-mapped arrays act almost exactly like regular in-memory NumPy arrays, so you can use all your usual array operations on the sliced portion without issues.
内容的提问来源于stack exchange,提问作者Justin

