Numba CUDA是否支持CUDA内置向量类型float3及实现方法
Numba CUDA实现float3等效功能方案
float3支持说明
Numba CUDA不直接提供CUDA C原生的float3内置向量类型,但有两种等效方案可以完全复现你给出的示例内核功能。
注:你提供的CUDA C示例存在变量名笔误,函数入参定义为
float3 *out,但内核中访问时用了未定义的dest变量,以下实现已修正该问题。
方案1:使用Numba内置向量类型模拟float3
Numba提供了Vec3f(3个单精度浮点数组成的向量),支持.x/.y/.z属性访问,使用逻辑和float3完全一致:
import numba from numba import cuda import numpy as np # 内核实现和CUDA C逻辑完全对齐 @cuda.jit def add_3darrs_broadcast(out, a, b, SZ): M = SZ[0] N = SZ[1] S = SZ[2] tx = cuda.threadIdx.x bx = cuda.blockIdx.x BSZ = cuda.blockDim.x for s in range(S): t = s * BSZ + tx if t < N: out[bx * N + t].x = b[t].x + a[bx].x out[bx * N + t].y = b[t].y + a[bx].y out[bx * N + t].z = b[t].z + a[bx].z cuda.syncthreads()
调用方法:
# 示例参数 M = 16 N = 1024 block_size = 256 grid_size = M SZ = np.array([M, N, N // block_size + 1], dtype=np.int32) # 构造Vec3f类型的数组 dtype_vec3f = numba.types.float3x3.as_dtype() a = np.array([(1.0, 2.0, 3.0) for _ in range(M)], dtype=dtype_vec3f) b = np.array([(4.0, 5.0, 6.0) for _ in range(N)], dtype=dtype_vec3f) out = np.zeros(M * N, dtype=dtype_vec3f) # 设备内存拷贝与内核启动 d_a = cuda.to_device(a) d_b = cuda.to_device(b) d_out = cuda.to_device(out) d_SZ = cuda.to_device(SZ) add_3darrs_broadcast[grid_size, block_size](d_out, d_a, d_b, d_SZ) # 取回结果 res = d_out.copy_to_host()
方案2:使用(N,3)格式二维数组实现(更推荐)
该方案不需要依赖特殊向量类型,直接用标准NumPy数组存储分量,使用更符合Python习惯,性能也更稳定:
@cuda.jit def add_3darrs_broadcast_v2(out, a, b, SZ): M = SZ[0] N = SZ[1] S = SZ[2] tx = cuda.threadIdx.x bx = cuda.blockIdx.x BSZ = cuda.blockDim.x for s in range(S): t = s * BSZ + tx if t < N: out[bx, t, 0] = b[t, 0] + a[bx, 0] out[bx, t, 1] = b[t, 1] + a[bx, 1] out[bx, t, 2] = b[t, 2] + a[bx, 2] cuda.syncthreads()
调用方法:
# 直接使用标准float32数组即可 a = np.full((M, 3), fill_value=[1,2,3], dtype=np.float32) b = np.full((N, 3), fill_value=[4,5,6], dtype=np.float32) out = np.zeros((M, N, 3), dtype=np.float32) # 后续设备拷贝、内核启动逻辑和方案1完全一致
内容的提问来源于stack exchange,提问作者andregalera
相关产品推荐
相关产品推荐

