You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Numpy向量化实现分箱值排序与多条件元素筛选求助

实现思路

核心是将三层优先级规则转化为可向量化计算的排序键,全程避免Python层面的逐行循环:

  1. 构造调整后的arr2值:对所有arr1元素为0的位置,将其对应的arr2值减去一个大于arr2全局极差的偏移量,保证所有arr1非零位置的调整后arr2值,一定大于所有arr1为零位置的调整后arr2值,既保留原始arr2的大小优先级,又自动实现「arr1非零优先」的规则。
  2. 每行按两级优先级降序选位置:第一优先级为调整后的arr2值(等价于从arr2最大值开始逐层向下找,直到找到存在非零arr1值的层级),第二优先级为原始arr1值(同一arr2值层级下选arr1最大的位置)。
  3. 调用np.lexsort沿行方向按上述优先级排序,取每行排序后的第一个位置即为目标索引,最后提取对应arr1值拼接为结果。

该实现所有逻辑均依赖Numpy底层C实现的向量化操作,处理50万行数据可达到百毫秒级性能。

代码实现
import numpy as np

def fancy_select(arr1: np.ndarray, arr2: np.ndarray) -> np.ndarray:
    # 计算偏移量,保证非零位置调整后arr2值一定大于零值位置
    arr2_min, arr2_max = arr2.min(), arr2.max()
    offset = (arr2_max - arr2_min) + 1
    adjusted_arr2 = arr2 - (arr1 == 0) * offset
    
    # lexsort优先级为传入键的逆序,传入负值实现降序效果,沿axis=1按行排序
    sorted_indices = np.lexsort((-arr1, -adjusted_arr2), axis=1)
    target_idx = sorted_indices[:, 0]
    
    # 提取目标值与索引拼接为结果
    row_indices = np.arange(arr1.shape[0])
    target_val = arr1[row_indices, target_idx]
    return np.stack([target_val, target_idx], axis=1)

更高性能版本(无溢出风险场景可用)

如果明确arr1、arr2的取值范围不会造成整数溢出,可直接计算复合得分替代lexsort,性能还可提升约20%:

def fancy_select_fast(arr1: np.ndarray, arr2: np.ndarray) -> np.ndarray:
    arr2_min, arr2_max = arr2.min(), arr2.max()
    offset = (arr2_max - arr2_min) + 1
    adjusted_arr2 = arr2 - (arr1 == 0) * offset
    
    # 给adjusted_arr2分配足够位权,保证优先级高于arr1
    score = adjusted_arr2 * (arr1.max() + 1) + arr1
    target_idx = score.argmax(axis=1)
    
    row_indices = np.arange(arr1.shape[0])
    target_val = arr1[row_indices, target_idx]
    return np.stack([target_val, target_idx], axis=1)
效果验证

用题目给出的示例测试:

arr1 = np.array([[0, 0, 4, 7, 3, 0, 0, 0, 0, 0],
                 [0, 3, 5, 7, 6, 0, 3, 0, 0, 0]])
arr2 = np.array([[14, 14, 14, 13, 11, 9, 6, 4, 2, 0],
                 [14, 13, 13, 13, 12, 9, 7, 4, 2, 0]])

print(fancy_select(arr1, arr2))
# 输出:
# [[4 2]
#  [7 3]]

与预期结果完全一致,同值候选返回任意一个均符合要求。

内容的提问来源于stack exchange,提问作者Gwalchaved

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 07:18:14