如何使用numpy清洗索引数组并同步调整关联值数组
实现代码
直接基于numpy原生方法即可完成需求,全程无循环,效率很高:
import numpy as np # 你的输入数据 delta = np.array([0,3,4,1,1,4,4,5,7,10], dtype = int) theta = np.random.normal(size = (12, 5)) # 步骤1:提取delta中所有出现过的索引,自动去重升序排列 used_indices = np.unique(delta) # 步骤2:生成清洗后的连续索引delta new_delta = np.searchsorted(used_indices, delta) # 步骤3:生成theta重排索引,匹配新的连续索引规则 all_indices = np.arange(theta.shape[0]) unused_indices = all_indices[~np.isin(all_indices, used_indices)] theta_reorder_idx = np.concatenate([used_indices, unused_indices]) new_theta = theta[theta_reorder_idx]
结果验证
运行后输出和给出的预期完全一致:
new_delta结果为array([0, 2, 3, 1, 1, 3, 3, 4, 5, 6])theta_reorder_idx结果为array([ 0, 1, 3, 4, 5, 7, 10, 2, 6, 8, 9, 11])
逻辑说明
np.unique提取所有出现过的索引,天然升序排列,正好对应需要的新连续索引的下标np.searchsorted利用二分查找直接映射原索引到新的连续位置,时间复杂度仅为O(n log k),k为唯一索引数量- 拼接已使用和未使用的索引得到theta重排规则,保证新索引值和theta的行位置一一对应
内容的提问来源于stack exchange,提问作者Faydey
相关产品推荐
相关产品推荐

