You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

寻求Numpy中高效实现自定义区间整数映射的函数

Efficient NumPy Alternative to Loop-Based Value Mapping

Great question—looping through NumPy arrays element-by-element is a common bottleneck, especially as your dataset scales. Let's replace that slow loop with vectorized NumPy operations that match your mapping logic perfectly.

Option 1: Use numpy.digitize() (Simplest & Fastest)

np.digitize() is built exactly for this kind of interval-based mapping. It takes your array and a set of bin edges, then returns the index of the bin each element falls into. Here's how to adapt it to your needs:

import numpy as np

def flatten(y):
    # Define your boundary points
    bins = np.array([0.7, 1.6, 2.4, 3.7])
    # Get bin indices (starts at 0 for values below the first bin)
    mapped_values = np.digitize(y, bins)
    
    # Quick breakdown of how this maps:
    # - Values < 0.7 → index 0 → matches your target 0
    # - 0.7 ≤ values < 1.6 → index 1 → matches your target 1
    # - 1.6 ≤ values < 2.4 → index 2 → matches your target 2
    # - 2.4 ≤ values < 3.7 → index 3 → matches your target 3
    # - Values ≥ 3.7 → index 4 → matches your target 4
    
    return mapped_values

Note: Your original loop uses strict inequalities (>, <) instead of ≥/≤ for boundaries. If you need to strictly exclude the boundary values (e.g., 0.7 should stay unmodified instead of mapping to 1), you can adjust by adding a quick mask to fix those points:

def flatten_strict(y):
    bins = np.array([0.7, 1.6, 2.4, 3.7])
    mapped_values = np.digitize(y, bins)
    # Reset values that equal the boundary points to their original value
    boundary_mask = np.isin(y, bins)
    mapped_values[boundary_mask] = y[boundary_mask]
    return mapped_values

Option 2: Use numpy.select() (Exact Loop Replication)

If you want to replicate your original loop's logic exactly (including all strict inequalities and unhandled boundary cases), np.select() lets you define explicit conditions and corresponding choices:

import numpy as np

def flatten(y):
    # Define your conditions in the same order as your loop
    conditions = [
        y < 0.7,
        (y > 0.7) & (y < 1.6),
        (y > 1.6) & (y < 2.4),
        (y > 2.4) & (y < 3.7),
        y > 3.7
    ]
    # Corresponding target values
    choices = [0, 1, 2, 3, 4]
    
    # Use default=y to keep original values where no condition is met (e.g., y=0.7)
    mapped_values = np.select(conditions, choices, default=y)
    return mapped_values

Why This Is Better Than Loops

NumPy vectorized operations run on C-level code instead of slow Python loops, so you'll see massive speedups—especially with large arrays. For example, an array of 1 million elements would take seconds with your loop, but milliseconds with either of these methods.

内容的提问来源于stack exchange,提问作者BorkoP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:43:10