寻求2D numpy数组指定元素删除后上移补位的高效实现方案
2D NumPy数组的平移补位实现优化方案
需求说明
给定一组(x, y)坐标(x为列索引,y为行索引),删除对应2D NumPy数组中的元素后,执行以下操作:
- 被删除元素所在列中,该元素下方的所有元素依次上移补位
- 数组底部的空缺位置用0-4之间的随机整数填充(示例中用
r标记)
示例展示
简单示例
输入
import numpy as np coordinates = [ (0, 1), (1, 1), (2, 1) ] test = np.array([ [0, 3, 0, 4, 0, 1, 3], [1, 1, 1, 0, 0, 4, 1], [2, 3, 3, 4, 1, 4, 3], [3, 1, 1, 3, 3, 2, 2], [2, 1, 3, 4, 3, 4, 4], [0, 0, 1, 1, 0, 2, 0], [0, 1, 0, 2, 3, 4, 2], [3, 2, 1, 1, 3, 2, 1], ])
输出
np.array([ [0, 3, 0, 4, 0, 1, 3], [2, 3, 3, 0, 0, 4, 1], [3, 1, 1, 4, 1, 4, 3], [2, 1, 3, 3, 3, 2, 2], [0, 0, 1, 4, 3, 4, 4], [0, 1, 0, 1, 0, 2, 0], [3, 2, 1, 2, 3, 4, 2], [r, r, r, 1, 3, 2, 1], ])
注:输出中的
r代表0-4的随机整数
复杂示例
输入
import numpy as np coordinates = [ (1, 2), (2, 2), (3, 0), (3, 1), (3, 2), (3, 3), (3, 4), (4, 2), (4, 3), (4, 4), (5, 2), (6, 2), ] test = np.array([ [0, 3, 0, 4, 1, 1, 3], [4, 1, 0, 4, 2, 4, 1], [2, 4, 4, 4, 3, 3, 3], [3, 1, 1, 4, 3, 2, 2], [2, 1, 3, 4, 3, 4, 4], [0, 0, 1, 1, 0, 2, 0], [0, 1, 0, 2, 3, 4, 2], [3, 2, 1, 1, 3, 2, 1], ])
输出
np.array([ [0, 3, 0, 1, 1, 1, 3], [4, 1, 0, 2, 2, 4, 1], [2, 1, 1, 1, 0, 2, 2], [3, 1, 3, r, 3, 4, 4], [2, 0, 1, r, 3, 2, 0], [0, 1, 0, r, r, 4, 2], [0, 2, 1, r, r, 2, 1], [3, r, r, r, r, r, r], ])
注:输出中的
r代表0-4的随机整数
现有实现
以下是能生成预期结果的基础代码,但存在三重循环,效率较低,尤其不适用于大规模数组:
import numpy as np NUM_ROWS = len(test) NUM_COLS = len(test[0]) # 从下到上遍历每一行 for row_i in range(NUM_ROWS - 1, -1, -1): for col_i in range(NUM_COLS): if (col_i, row_i) in coordinates: # 该列元素上移 for lower_row_i in range(row_i + 1, NUM_ROWS): test[lower_row_i - 1][col_i] = test[lower_row_i][col_i] # 底部填充随机值 test[NUM_ROWS - 1][col_i] = np.random.randint(5) print(test)
优化实现方案
方案1:按列向量化处理(高效推荐)
利用NumPy的向量化操作替代嵌套循环,大幅提升效率,尤其适合大数组:
import numpy as np def shift_and_fill(test, coordinates): rows, cols = test.shape # 将坐标转换为numpy数组,方便分组处理 coords = np.array(coordinates) # 按列分组,获取每列需要删除的行索引 for col in np.unique(coords[:, 0]): # 获取当前列需要删除的所有行索引,降序排列(从下往上处理) del_rows = coords[coords[:, 0] == col][:, 1] del_rows = np.sort(del_rows)[::-1] for row in del_rows: # 该列从row+1到末尾的元素上移一位 test[row:-1, col] = test[row+1:, col] # 底部填充随机值 test[-1, col] = np.random.randint(5) return test # 使用示例 result = shift_and_fill(test.copy(), coordinates) print(result)
优势:
- 减少循环层级,仅保留列循环和列内删除行的循环
- 利用NumPy切片操作替代内层循环,提升运算效率
- 逻辑清晰,便于维护
方案2:掩码批量处理(更简洁)
通过创建掩码标记需要保留的元素,再对每列重新构造数组:
import numpy as np def mask_based_shift(test, coordinates): rows, cols = test.shape # 创建全True的掩码,标记需要保留的元素 mask = np.ones_like(test, dtype=bool) # 将需要删除的坐标标记为False x_coords, y_coords = zip(*coordinates) mask[y_coords, x_coords] = False for col in range(cols): # 获取当前列需要保留的元素 kept = test[:, col][mask[:, col]] # 计算需要补充的随机值数量 fill_num = rows - len(kept) # 构造新列:保留元素 + 随机值 new_col = np.concatenate([kept, np.random.randint(0,5, fill_num)]) test[:, col] = new_col return test # 使用示例 result = mask_based_shift(test.copy(), coordinates) print(result)
优势:
- 代码更简洁,逻辑直观
- 批量处理每列,避免逐行判断
- 适合需要明确区分保留/删除元素的场景
方案3:循环优化(兼容小规模数组)
对原代码进行小幅优化,减少不必要的查询操作:
import numpy as np def optimized_loop(test, coordinates): rows, cols = test.shape # 将坐标转换为集合,提升查询速度(O(1)查询) coord_set = set(coordinates) # 从下到上遍历行 for row_i in range(rows-1, -1, -1): for col_i in range(cols): if (col_i, row_i) in coord_set: # 切片操作替代内层循环 test[row_i:-1, col_i] = test[row_i+1:, col_i] test[-1, col_i] = np.random.randint(5) return test # 使用示例 result = optimized_loop(test.copy(), coordinates) print(result)
优势:
- 仅修改原代码的关键部分,学习成本低
- 将坐标列表转为集合,大幅提升
in操作的查询速度 - 用切片替代内层循环,提升执行效率
内容的提问来源于stack exchange,提问作者zccafa3
相关产品推荐
相关产品推荐

