You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在NumPy/CuPy中按列分别不同偏移量滚动大型数组?

每列不同偏移量的数组滚动实现(GPU优先)

最优GPU方案(CuPy)

利用CuPy的广播索引实现全向量化操作,完全适配GPU并行计算,是大型数组场景下的最快解决方案。核心思路是构造索引矩阵直接定位每个输出元素的原数组位置,避免逐列循环的开销。

import cupy as cp

def roll_per_column_gpu(x):
    n_rows, n_cols = x.shape
    # 定义每列的偏移量:第j列(0-based)偏移j+1位
    offsets = cp.arange(1, n_cols + 1)
    # 生成广播后的行索引:每行索引减去对应列偏移量,取模行数实现循环滚动
    row_indices = (cp.arange(n_rows)[:, None] - offsets) % n_rows
    # GPU并行完成元素提取
    return x[row_indices, cp.arange(n_cols)]

测试验证

x = cp.reshape(cp.arange(15), (5, 3))
y = roll_per_column_gpu(x)
print(cp.asnumpy(y))
# 输出结果:
# [[12 10  8]
#  [ 0 13 11]
#  [ 3  1 14]
#  [ 6  4  2]
#  [ 9  7  5]]

CPU方案(NumPy)

思路与GPU版本完全一致,利用NumPy的广播索引实现向量化计算,适合无法使用GPU的场景:

import numpy as np

def roll_per_column_cpu(x):
    n_rows, n_cols = x.shape
    offsets = np.arange(1, n_cols + 1)
    row_indices = (np.arange(n_rows)[:, None] - offsets) % n_rows
    return x[row_indices, np.arange(n_cols)]

方案说明

  • 通用性:若需自定义每列偏移量,仅需修改offsets数组即可(例如offsets = cp.array([2, 4, 1])对应列1偏移2位、列2偏移4位等)。
  • 性能优势:向量化操作避免了Python循环的额外开销,CuPy版本在GPU上自动调度并行线程,处理百万级以上大型数组时,性能远超逐列调用cp.roll的循环方案。

内容的提问来源于stack exchange,提问作者Jeffrey Chen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 06:01:44