如何高效实现支持整数与半值索引的Numpy数组?
Instead of subclassing numpy.ndarray (which leads to conflicts with numpy's internal index handling), use a composition-based wrapper class to control index conversion explicitly. This approach preserves numpy's native indexing features (slicing, boolean masks, etc.) while adding support for integer and half-value indices, with minimal performance overhead.
Implementation Code
import numpy as np class HalfIndexedArray: def __init__(self, data): self.data = np.asarray(data) def _process_index(self, idx): # Convert integer/half-integer indices to underlying array indices if isinstance(idx, (float, np.floating)): doubled = idx * 2 if not np.isclose(doubled, round(doubled)): raise ValueError(f"Index {idx} must be integer or half-integer") return int(round(doubled)) elif isinstance(idx, (int, np.integer)): return idx * 2 elif isinstance(idx, slice): # Process slice bounds if they are integer/half-integer start = self._process_index(idx.start) if idx.start is not None else None stop = self._process_index(idx.stop) if idx.stop is not None else None step = self._process_index(idx.step) if idx.step is not None else None return slice(start, stop, step) elif isinstance(idx, tuple): return tuple(self._process_index(i) for i in idx) elif idx is Ellipsis: return Ellipsis else: # Pass through other index types (boolean arrays, integer arrays) return idx def __getitem__(self, key): processed_key = self._process_index(key) result = self.data[processed_key] # Wrap array results to maintain half-indexing functionality return HalfIndexedArray(result) if isinstance(result, np.ndarray) else result def __setitem__(self, key, value): processed_key = self._process_index(key) self.data[processed_key] = value # Expose common numpy array attributes @property def shape(self): return self.data.shape @property def dtype(self): return self.data.dtype @property def ndim(self): return self.data.ndim @property def size(self): return self.data.size # Allow conversion to numpy array for direct operations def __array__(self, dtype=None): return self.data.astype(dtype) if dtype else self.data # Support numpy ufuncs (arithmetic, math operations) def __array_ufunc__(self, ufunc, method, *inputs, **kwargs): processed_inputs = [inp.data if isinstance(inp, HalfIndexedArray) else inp for inp in inputs] result = ufunc(*processed_inputs, **kwargs) return HalfIndexedArray(result) if isinstance(result, np.ndarray) else result
Key Features
Index Conversion:
- Integer indices
kmap to underlying index2*k - Half-integer indices
k+0.5map to underlying index2*(k+0.5) = 2k+1 - Validates that float indices are either integer or half-integer to avoid invalid access
- Integer indices
Preserves Numpy Functionality:
- Full support for slicing (e.g.,
a[0:1.5]maps to underlying slice0:3) - Works with boolean masks, integer array indices, and ellipsis (
...) - Numpy ufuncs (e.g.,
np.sum,a + 10) operate directly on the underlying array for performance
- Full support for slicing (e.g.,
Performance:
- Minimal overhead from index processing (O(1) per index component)
- All heavy computations use numpy's optimized operations on the underlying array
Example Usage
# Initialize the array a = HalfIndexedArray([[0,1,2,3],[10,11,12,13],[20,21,22,23]]) # Basic indexing print(a[0,0]) # Output: 0 print(a[0.5,0.5]) # Output:11 print(a[0,1+0.5]) # Output:3 # Slicing print(a[0:1].data) # Output: [[0,1,2,3],[10,11,12,13]] print(a[0.5:1].data) # Output: [[10,11,12,13]] # Modify values a[0.5,0.5] = 99 print(a[0.5,0.5]) # Output:99 # Numpy operations b = a * 2 print(b[0.5,0.5]) # Output:198
Why This Fixes the Subclassing Issue
Subclassing numpy.ndarray leads to conflicts because numpy's internal index handling may reapply your conversion logic multiple times (e.g., when returning views). The composition approach keeps full control over index processing, ensuring each index is converted exactly once before accessing the underlying array.
内容的提问来源于stack exchange,提问作者D__

