逻辑数组指定索引范围的最近前置True索引映射方案优化
Great question! When working with large arrays (think millions of elements), your original loop-based approach will grind to a halt because of the repeated np.where calls inside the loop—each one is an O(n) operation, and doing that in a loop scales terribly. Let's replace it with a fully vectorized numpy solution that's way faster, cleaner, and easier to read.
What's Wrong with the Original Code?
The bottleneck is this line inside the loop:
index_loc = np.where(true_indices <= index)[0][-1]
For every element in irange, you're scanning the entire true_indices array to find elements ≤ current index. That's an O(k*m) time complexity (k = length of irange, m = number of True indices), which is brutal for large datasets.
The Fix: Use np.searchsorted
Since true_indices is naturally sorted (we get it from np.where, which returns indices in ascending order), we can use numpy.searchsorted to do binary searches for each element in irange—this cuts the time complexity to O(k*log(m)), which is night-and-day faster for big data.
Optimized Code
import numpy as np def map_nearest_preceding_true_indices(tmask, irange): true_indices = np.where(tmask)[0] # Find the insertion point for each index in irange (right side) # Subtract 1 to get the last True index ≤ current index positions = np.searchsorted(true_indices, irange, side='right') - 1 # Vectorized lookup—no loops needed! return true_indices[positions]
Why This Works (and Is Faster)
np.searchsorteduses binary search under the hood, which is way more efficient than linear scans for sorted arrays- Everything is vectorized—no Python-level loops, so we leverage numpy's optimized C backend for maximum speed
Test It With Your Example
Let's verify with your sample input:
tmask = np.array([1, 0, 0, 1, 1, 1, 0, 1, 0, 1, 0], dtype=bool) irange = np.array([0, 1, 2, 3, 4, 5, 6, 7]) print(map_nearest_preceding_true_indices(tmask, irange)) # Output: [0 0 0 3 4 5 5 7] — exactly what you expected!
Handling Edge Cases
What if an index in irange is smaller than the first True index in tmask? The code above would return the last element of true_indices (since positions would be -1, and numpy wraps around negative indices). To handle this, add a quick check:
def map_nearest_preceding_true_indices(tmask, irange): true_indices = np.where(tmask)[0] if not len(true_indices): raise ValueError("Input mask has no True values—can't find preceding indices!") positions = np.searchsorted(true_indices, irange, side='right') - 1 # Set positions < 0 to 0 (or another default like -1 if you prefer) positions[positions < 0] = 0 return true_indices[positions]
This ensures you get a valid index even for values before the first True.
内容的提问来源于stack exchange,提问作者Wall-E

