You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Numpy高效实现:在数组中搜索另一数组元素的优化方案

高效实现Numpy数组元素匹配索引需求

我有两个Numpy数组:

first_array = np.array([2,2,2,10,10,15,20,20,20,20])
second_array = np.array([15,5,10,78,2,44,20,2,66,1,10,15,40,85,71,23,20,45,29,20,1])

需求是:将first_array中的每个元素在second_array中搜索,最终得到一个二维数组,包含second_array的索引和对应搜索值。

目前使用的循环代码可实现需求,但效率偏低,希望找到无需逐元素循环的高效实现方法:

out = []
for i in first_array:
    index = np.argwhere(second_array == i)
    out.append(np.array([*index.T,np.ones(len(index))*i]))
    
np.hstack(out).T

预期输出:

desired_output = np.array([[4,7,4,7,4,7,2,10,2,10,0,11,6,16,19,6,16,19,6,16,19,6,16,19],
                         [2,2,2,2,2,2,10,10,10,10,15,15,20,20,20,20,20,20,20,20,20,20,20,20]]).T

高效解决方案

可以利用Numpy的向量化操作避免循环,大幅提升效率,以下是两种实现方式:

方式一:基于值-索引映射

先建立second_array中值到对应索引的映射,避免重复搜索同一值,再根据first_array的元素批量生成结果:

import numpy as np

first_array = np.array([2,2,2,10,10,15,20,20,20,20])
second_array = np.array([15,5,10,78,2,44,20,2,66,1,10,15,40,85,71,23,20,45,29,20,1])

# 建立值到索引列表的映射
val_to_indices = {val: np.where(second_array == val)[0] for val in np.unique(second_array)}

# 批量获取每个元素对应的索引和值
indices_list = []
values_list = []
for val in first_array:
    indices = val_to_indices[val]
    indices_list.append(indices)
    values_list.append(np.full_like(indices, val))

# 拼接成最终二维数组
result = np.column_stack([np.concatenate(indices_list), np.concatenate(values_list)])

方式二:基于重复次数统计

先统计first_array中各元素的出现次数,直接将对应索引数组重复对应次数,进一步减少计算量:

import numpy as np

first_array = np.array([2,2,2,10,10,15,20,20,20,20])
second_array = np.array([15,5,10,78,2,44,20,2,66,1,10,15,40,85,71,23,20,45,29,20,1])

# 统计first_array中各唯一值的出现次数
unique_vals, repeat_counts = np.unique(first_array, return_counts=True)

# 生成重复后的索引和值数组
repeated_indices = []
repeated_vals = []
for val, cnt in zip(unique_vals, repeat_counts):
    indices = np.where(second_array == val)[0]
    repeated_indices.append(np.tile(indices, cnt))  # 将索引数组重复cnt次
    repeated_vals.append(np.full(len(indices)*cnt, val))  # 生成对应次数的数值数组

# 拼接结果
result = np.column_stack([np.concatenate(repeated_indices), np.concatenate(repeated_vals)])

两种方式生成的result均与desired_output完全一致,且相比原循环代码,避免了重复搜索同一元素,在数组规模较大时效率提升显著。

内容的提问来源于stack exchange,提问作者datatech

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 11:55:20