寻求Numpy中高效实现自定义区间整数映射的函数
Great question—looping through NumPy arrays element-by-element is a common bottleneck, especially as your dataset scales. Let's replace that slow loop with vectorized NumPy operations that match your mapping logic perfectly.
Option 1: Use numpy.digitize() (Simplest & Fastest)
np.digitize() is built exactly for this kind of interval-based mapping. It takes your array and a set of bin edges, then returns the index of the bin each element falls into. Here's how to adapt it to your needs:
import numpy as np def flatten(y): # Define your boundary points bins = np.array([0.7, 1.6, 2.4, 3.7]) # Get bin indices (starts at 0 for values below the first bin) mapped_values = np.digitize(y, bins) # Quick breakdown of how this maps: # - Values < 0.7 → index 0 → matches your target 0 # - 0.7 ≤ values < 1.6 → index 1 → matches your target 1 # - 1.6 ≤ values < 2.4 → index 2 → matches your target 2 # - 2.4 ≤ values < 3.7 → index 3 → matches your target 3 # - Values ≥ 3.7 → index 4 → matches your target 4 return mapped_values
Note: Your original loop uses strict inequalities (>, <) instead of ≥/≤ for boundaries. If you need to strictly exclude the boundary values (e.g., 0.7 should stay unmodified instead of mapping to 1), you can adjust by adding a quick mask to fix those points:
def flatten_strict(y): bins = np.array([0.7, 1.6, 2.4, 3.7]) mapped_values = np.digitize(y, bins) # Reset values that equal the boundary points to their original value boundary_mask = np.isin(y, bins) mapped_values[boundary_mask] = y[boundary_mask] return mapped_values
Option 2: Use numpy.select() (Exact Loop Replication)
If you want to replicate your original loop's logic exactly (including all strict inequalities and unhandled boundary cases), np.select() lets you define explicit conditions and corresponding choices:
import numpy as np def flatten(y): # Define your conditions in the same order as your loop conditions = [ y < 0.7, (y > 0.7) & (y < 1.6), (y > 1.6) & (y < 2.4), (y > 2.4) & (y < 3.7), y > 3.7 ] # Corresponding target values choices = [0, 1, 2, 3, 4] # Use default=y to keep original values where no condition is met (e.g., y=0.7) mapped_values = np.select(conditions, choices, default=y) return mapped_values
Why This Is Better Than Loops
NumPy vectorized operations run on C-level code instead of slow Python loops, so you'll see massive speedups—especially with large arrays. For example, an array of 1 million elements would take seconds with your loop, but milliseconds with either of these methods.
内容的提问来源于stack exchange,提问作者BorkoP

