You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

numpy.memmap操作能否延迟至访问时?附实例分析

Lazy Operations for NumPy Memmap: Delay Calculations Until Access

Great question—you’re spot-on that subclassing is the way to go here. The core issue is that NumPy’s default ufuncs (like +, *) return regular ndarray instances when applied to memmap, which forces the entire file to load into memory. To fix this, we can build a lazy wrapper that queues operations instead of executing them immediately, only computing values when you actually access elements.

A Lazy Memmap Implementation

Here’s a practical subclass of ndarray that holds a memmap and a chain of pending operations. It only computes values when you access specific elements or slices:

import numpy as np

class LazyArray(np.ndarray):
    def __new__(cls, memmap_obj, operations=None):
        # Initialize the array shell without loading data
        obj = super().__new__(cls, shape=memmap_obj.shape, dtype=memmap_obj.dtype)
        obj.memmap = memmap_obj
        obj.operations = operations or []
        return obj

    def __array_ufunc__(self, ufunc, method, *inputs, **kwargs):
        # Queue the operation instead of executing it right away
        new_ops = self.operations.copy()
        # Process inputs to handle nested LazyArrays
        processed_inputs = []
        for inp in inputs:
            if isinstance(inp, LazyArray):
                # Store the underlying memmap and its operation chain
                processed_inputs.append((inp.memmap, inp.operations))
            else:
                processed_inputs.append(inp)
        new_ops.append((ufunc, method, processed_inputs, kwargs))
        # Return a new LazyArray with the updated operation queue
        return LazyArray(self.memmap, new_ops)

    def __getitem__(self, key):
        # Compute results only when elements are accessed
        result = self.memmap[key]
        # Apply all pending operations to the accessed slice/element
        for ufunc, _, inputs, kwargs in self.operations:
            processed_inps = []
            for inp in inputs:
                if isinstance(inp, tuple):
                    # Resolve nested LazyArray operations
                    sub_memmap, sub_ops = inp
                    sub_result = sub_memmap[key]
                    for sub_ufunc, _, sub_inps, sub_kwargs in sub_ops:
                        sub_result = sub_ufunc(sub_result, **sub_kwargs)
                    processed_inps.append(sub_result)
                else:
                    processed_inps.append(inp)
            result = ufunc(*processed_inps, **kwargs)
        return result

    def __repr__(self):
        # Avoid triggering full array load in string representation
        return f"LazyArray(shape={self.shape}, dtype={self.dtype}, pending_ops={len(self.operations)})"

How to Use It

Let’s test this with your original example:

# Create a large test array and save it
a = np.arange(10_000_000, dtype=np.int32)
np.save("a.npy", a)

# Load with memmap and wrap in LazyArray
mmap_a = np.load("a.npy", mmap_mode='r')
lazy_a = LazyArray(mmap_a)

# Perform operations—no data is loaded yet!
lazy_b = lazy_a + 2
lazy_c = lazy_b * 3

# Check the lazy state
print(lazy_c)  # Output: LazyArray(shape=(10000000,), dtype=int32, pending_ops=2)

# Access an element—only this position is loaded and computed
print(lazy_c[100])  # Output: 306 (since (100 + 2) * 3 = 306)

# Convert to regular ndarray if you need full in-memory access
full_c = np.array(lazy_c)
print(type(full_c))  # Output: <class 'numpy.ndarray'>

Key Notes

  • Element-Level Only: This works best for element-wise operations. Functions that require global data (like sum(), mean()) will still trigger a full load, since they need to access every element.
  • Slice Support: The __getitem__ method handles slices too—if you access lazy_c[100:200], only that slice is loaded and computed.
  • Disk Efficiency: Since we’re still using memmap, data is read from disk in blocks only when needed, keeping memory usage low.

内容的提问来源于stack exchange,提问作者bers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:56:37