通过numpy memmap读取的值求和结果异常问题
np.memmap Returns Incorrect Sum Compared to np.load for .npy Files The Root Cause
The key issue here is that .npy files aren't raw binary arrays—they have a structured header at the start that stores metadata (like dtype, shape, endianness, etc.). When you use np.memmap directly with the full array shape and uint8 dtype, you're reading the header bytes as part of your array data, which shifts all subsequent values and leads to incorrect calculations.
Your original memmap call treats the entire file as raw uint8 data, but the first few hundred bytes are actually the npy header, not your array's first elements. That's why the sum differs from np.load, which properly parses the header and reads only the actual array data.
Fixes to Correctly Memmap the .npy File
You have two reliable ways to fix this:
1. Use np.load with mmap_mode (Simplest & Recommended)
Let numpy handle the header parsing automatically by using np.load with the mmap_mode parameter. This creates a memory-mapped array without loading the entire file into RAM, while respecting the npy file structure:
x = np.load('file.npy', mmap_mode='r') print(np.sum(x[0,0,:,0])) # Should match the np.load result
2. Manually Skip the NPY Header (Advanced)
If you need to use np.memmap directly, you first need to calculate the length of the npy header to set the correct offset parameter. Here's how to do it:
import numpy as np # Open the file and read the header to get its total length with open('file.npy', 'rb') as f: # Validate magic bytes for npy format magic = f.read(6) if magic != b'\x93NUMPY': raise ValueError("Not a valid .npy file") # Read version number version = f.read(2) # Read header length (4 bytes for version >= 2.0, 2 bytes for older) if version >= b'\x02\x00': header_len = np.frombuffer(f.read(4), dtype=np.uint32)[0] else: header_len = np.frombuffer(f.read(2), dtype=np.uint16)[0] # Calculate total bytes occupied by the header total_header_len = 6 + 2 + (4 if version >= b'\x02\x00' else 2) + header_len # Create memmap with offset to skip the header x = np.memmap('file.npy', dtype='uint8', shape=(67000, 9, 128, 128), mode='r', offset=total_header_len) print(np.sum(x[0,0,:,0])) # Should now match the np.load result
Why This Happens in Detail
When you save an array with np.save, numpy writes four distinct parts to the file:
- A 6-byte magic string:
\x93NUMPY - A 2-byte version number (e.g.,
\x01\x00for version 1.0) - A header length indicator (2 or 4 bytes, depending on the version)
- A human-readable JSON header with metadata (like
{"descr": "<u1", "shape": [67000,9,128,128], ...})
Your original memmap call reads all these bytes as part of the array, so the first element of your memmap isn't the first element of your original array—it's the first byte of the magic string (\x93, which is 147 in uint8), throwing off all subsequent calculations.
内容的提问来源于stack exchange,提问作者Samuel Prevost

