NumPy内存映射随机切片效率问题咨询
Let me break down exactly why you're seeing this performance gap, and how to fix it.
The Root Cause: Memory Layout & Disk Storage
NumPy uses C-style (row-major) ordering by default for arrays, including memory-mapped ones. This means your 100k×100k float16 array is stored on disk row-by-row:
- All elements of row 0 come first, then row 1, then row 2, and so on.
- Each float16 is 2 bytes, so a single row takes up
100000 * 2 = 200KBof contiguous disk space.
When you take a vertical slice like fp_read[:,17000], you're asking for the 17000th element from every row. On disk, these elements are not contiguous: each is separated by 200KB (the size of an entire row). Reading this requires 100,000 separate random disk reads—each fetching just 2 bytes. Disk hardware is terrible at random IO compared to contiguous reads, hence the poor performance.
In contrast, horizontal slices (like fp_read[50000,:]) or even block slices (like fp_read[1000:2000, 1000:2000]) pull contiguous chunks of data, which the OS can prefetch in large blocks, making them orders of magnitude faster.
Fixes to Speed Up Vertical Slices
Here are two actionable solutions depending on whether you can re-generate the data:
1. Re-save the Data in Fortran (Column-Major) Order
If you have access to the original data or can re-generate the memmap file, save it using Fortran-style (column-major) ordering. This stores columns contiguously on disk, making vertical slices as fast as horizontal slices were before.
# Create a new memmap file with Fortran order fp_write = np.memmap('column_ordered_data.dat', dtype='float16', mode='w+', shape=(100000, 100000), order='F') # Copy your existing data into it (or write new data directly) fp_write[:] = original_data[:] # Clean up to ensure data is flushed to disk del fp_write # Now read the column-ordered file fp_read_col = np.memmap('column_ordered_data.dat', dtype='float16', mode='r', shape=(100000, 100000), order='F') # This vertical slice will now be fast! %timeit fp_read_col[:,17000]
2. Batch Reads for Existing C-Order Data
If you can't re-save the file, optimize your reads to minimize random IO:
- Instead of reading a single column at a time, read large blocks of rows first, then extract the columns you need from those rows. This turns random reads into contiguous reads.
- Example:
# Read 1000 rows at a time, then extract column 17000 from each block def read_column_in_blocks(col_idx, block_size=1000): result = [] for i in range(0, 100000, block_size): block = fp_read[i:i+block_size, :] result.append(block[:, col_idx]) return np.concatenate(result) # This will be much faster than reading the column directly %timeit read_column_in_blocks(17000)
Key Takeaway
Always match your array's storage order to the type of slices you'll be accessing most. For column-heavy operations, Fortran order is your friend; for row-heavy work, stick with default C order. Disk IO patterns make all the difference here!
内容的提问来源于stack exchange,提问作者Attack68

