如何从含行列索引与值的列表高效构建numpy数组?
高效构建指定索引赋值的NumPy数组
给定列表l,每个元素格式为[row_index, column_index, value],需要将对应索引位置的NumPy数组元素设为指定值。替代低效的循环填充,以下是两种高效实现方式:
方法一:矢量化索引批量赋值
利用NumPy的矢量化操作直接批量处理索引和值,完全避免Python循环:
import numpy as np l = [[0, 0, 0.0], [0, 1, 14.0], [0, 2, 7.0], [1, 1, 0.0], [1, 2, 7.0], [2, 2, 0.0], [1, 0, 14.0], [2, 0, 7.0], [2, 1, 7.0]] # 将列表转为NumPy数组,拆分索引与值 arr = np.array(l) rows = arr[:, 0].astype(int) cols = arr[:, 1].astype(int) values = arr[:, 2] # 确定目标数组的形状 max_row, max_col = rows.max(), cols.max() array_ = np.zeros((max_row + 1, max_col + 1), dtype=values.dtype) # 批量赋值(矢量化操作,远快于循环) array_[rows, cols] = values print(array_)
输出结果:
[[ 0. 14. 7.] [14. 0. 7.] [ 7. 7. 0.]]
方法二:稀疏矩阵转稠密数组(适合稀疏场景)
如果列表中仅少数索引位置有非零值,使用稀疏矩阵构建再转换为稠密数组,内存效率更高:
import numpy as np from scipy.sparse import coo_matrix l = [[0, 0, 0.0], [0, 1, 14.0], [0, 2, 7.0], [1, 1, 0.0], [1, 2, 7.0], [2, 2, 0.0], [1, 0, 14.0], [2, 0, 7.0], [2, 1, 7.0]] arr = np.array(l) rows = arr[:, 0].astype(int) cols = arr[:, 1].astype(int) values = arr[:, 2] max_row, max_col = rows.max(), cols.max() # 构建COO格式稀疏矩阵 coo_array = coo_matrix((values, (rows, cols)), shape=(max_row + 1, max_col + 1)) # 转换为稠密NumPy数组 array_ = coo_array.toarray() print(array_)
核心优势
两种方法均使用NumPy/scipy的底层矢量化实现,避开了Python循环的性能瓶颈,数据量越大,效率提升越显著。
内容的提问来源于stack exchange,提问作者Akira
相关产品推荐
相关产品推荐

