如何向量化快速从NumPy一维数组提取变长切片并插入二维数组
从NumPy一维数组提取变长切片并注入二维数组的高效向量化方法
需求说明
需从单个NumPy一维数组中提取多个不重叠的变长切片,插入到二维数组的对应行中。
示例数据
- 源一维数组:
import numpy as np s = np.array([0, 1, 2, 3, 4, 5, 6, 7, 9, 10, 11, 12])
- 目标二维零数组:
t = np.array([[0, 0, 0, 0, 0], [0, 0, 0, 0, 0]])
- 提取与注入要求:取
s[1:4:1]和s[6:10:1]分别注入t的第0行和第1行,期望结果:
array([[1, 2, 3, 0, 0], [6, 7, 9, 10, 0]])
现有循环实现及问题
等长切片可通过二维索引直接实现,但变长切片无法直接套用。当前使用循环完成需求,但数据量较大时效率不足:
source = np.array([0, 1, 2, 3, 4, 5, 6, 7, 9, 10, 11, 12]) source_slices = [np.s_[0:2:1], np.s_[7:10:1]] target_slices = [np.s_[0:2:1], np.s_[0:3:1]] target = np.zeros(shape=(2,10), dtype=np.int32) # 循环赋值逻辑 for index in range(len(source_slices)): target[index, target_slices[index]] = source[source_slices[index]]
执行后结果:
[[ 0 1 0 0 0 0 0 0 0 0] [ 7 9 10 0 0 0 0 0 0 0]]
向量化高效实现方案
核心思路是将所有切片的索引扁平化,通过一次性赋值替代循环,完全利用NumPy的向量化运算优势。
代码实现
import numpy as np source = np.array([0, 1, 2, 3, 4, 5, 6, 7, 9, 10, 11, 12]) # 定义源切片(起始索引+长度)和目标切片(行索引+起始列+长度) source_specs = [(1, 3), (6, 4)] target_specs = [(0, 0, 3), (1, 0, 4)] target = np.zeros(shape=(2, 5), dtype=np.int32) # 生成源数组中需要提取的所有元素索引 source_indices = np.concatenate([np.arange(start, start + length) for start, length in source_specs]) # 生成目标数组对应位置的扁平化索引(将二维坐标转为一维) target_indices = np.concatenate([ np.ravel_multi_index((row, np.arange(col_start, col_start + length)), target.shape) for row, col_start, length in target_specs ]) # 一次性完成赋值 target.flat[target_indices] = source[source_indices] print(target)
输出结果:
[[1 2 3 0 0] [6 7 9 10 0]]
方案优势
- 完全规避Python循环,数据量越大,效率提升越显著
- 逻辑统一,通过扁平化索引处理所有变长切片的赋值需求
- 支持任意数量的变长切片,扩展性强
内容的提问来源于stack exchange,提问作者scotsman60
相关产品推荐
相关产品推荐

