numpy.memmap操作能否延迟至访问时?附实例分析
Lazy Operations for NumPy Memmap: Delay Calculations Until Access
Great question—you’re spot-on that subclassing is the way to go here. The core issue is that NumPy’s default ufuncs (like +, *) return regular ndarray instances when applied to memmap, which forces the entire file to load into memory. To fix this, we can build a lazy wrapper that queues operations instead of executing them immediately, only computing values when you actually access elements.
A Lazy Memmap Implementation
Here’s a practical subclass of ndarray that holds a memmap and a chain of pending operations. It only computes values when you access specific elements or slices:
import numpy as np class LazyArray(np.ndarray): def __new__(cls, memmap_obj, operations=None): # Initialize the array shell without loading data obj = super().__new__(cls, shape=memmap_obj.shape, dtype=memmap_obj.dtype) obj.memmap = memmap_obj obj.operations = operations or [] return obj def __array_ufunc__(self, ufunc, method, *inputs, **kwargs): # Queue the operation instead of executing it right away new_ops = self.operations.copy() # Process inputs to handle nested LazyArrays processed_inputs = [] for inp in inputs: if isinstance(inp, LazyArray): # Store the underlying memmap and its operation chain processed_inputs.append((inp.memmap, inp.operations)) else: processed_inputs.append(inp) new_ops.append((ufunc, method, processed_inputs, kwargs)) # Return a new LazyArray with the updated operation queue return LazyArray(self.memmap, new_ops) def __getitem__(self, key): # Compute results only when elements are accessed result = self.memmap[key] # Apply all pending operations to the accessed slice/element for ufunc, _, inputs, kwargs in self.operations: processed_inps = [] for inp in inputs: if isinstance(inp, tuple): # Resolve nested LazyArray operations sub_memmap, sub_ops = inp sub_result = sub_memmap[key] for sub_ufunc, _, sub_inps, sub_kwargs in sub_ops: sub_result = sub_ufunc(sub_result, **sub_kwargs) processed_inps.append(sub_result) else: processed_inps.append(inp) result = ufunc(*processed_inps, **kwargs) return result def __repr__(self): # Avoid triggering full array load in string representation return f"LazyArray(shape={self.shape}, dtype={self.dtype}, pending_ops={len(self.operations)})"
How to Use It
Let’s test this with your original example:
# Create a large test array and save it a = np.arange(10_000_000, dtype=np.int32) np.save("a.npy", a) # Load with memmap and wrap in LazyArray mmap_a = np.load("a.npy", mmap_mode='r') lazy_a = LazyArray(mmap_a) # Perform operations—no data is loaded yet! lazy_b = lazy_a + 2 lazy_c = lazy_b * 3 # Check the lazy state print(lazy_c) # Output: LazyArray(shape=(10000000,), dtype=int32, pending_ops=2) # Access an element—only this position is loaded and computed print(lazy_c[100]) # Output: 306 (since (100 + 2) * 3 = 306) # Convert to regular ndarray if you need full in-memory access full_c = np.array(lazy_c) print(type(full_c)) # Output: <class 'numpy.ndarray'>
Key Notes
- Element-Level Only: This works best for element-wise operations. Functions that require global data (like
sum(),mean()) will still trigger a full load, since they need to access every element. - Slice Support: The
__getitem__method handles slices too—if you accesslazy_c[100:200], only that slice is loaded and computed. - Disk Efficiency: Since we’re still using
memmap, data is read from disk in blocks only when needed, keeping memory usage low.
内容的提问来源于stack exchange,提问作者bers
相关产品推荐
相关产品推荐

