You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Numba CUDA是否支持CUDA内置向量类型float3及实现方法

Numba CUDA实现float3等效功能方案

float3支持说明

Numba CUDA不直接提供CUDA C原生的float3内置向量类型,但有两种等效方案可以完全复现你给出的示例内核功能。

注:你提供的CUDA C示例存在变量名笔误,函数入参定义为float3 *out,但内核中访问时用了未定义的dest变量,以下实现已修正该问题。

方案1:使用Numba内置向量类型模拟float3

Numba提供了Vec3f(3个单精度浮点数组成的向量),支持.x/.y/.z属性访问,使用逻辑和float3完全一致:

import numba
from numba import cuda
import numpy as np

# 内核实现和CUDA C逻辑完全对齐
@cuda.jit
def add_3darrs_broadcast(out, a, b, SZ):
    M = SZ[0]
    N = SZ[1]
    S = SZ[2]
    tx = cuda.threadIdx.x
    bx = cuda.blockIdx.x
    BSZ = cuda.blockDim.x

    for s in range(S):
        t = s * BSZ + tx
        if t < N:
            out[bx * N + t].x = b[t].x + a[bx].x
            out[bx * N + t].y = b[t].y + a[bx].y
            out[bx * N + t].z = b[t].z + a[bx].z
        cuda.syncthreads()

调用方法:

# 示例参数
M = 16
N = 1024
block_size = 256
grid_size = M
SZ = np.array([M, N, N // block_size + 1], dtype=np.int32)

# 构造Vec3f类型的数组
dtype_vec3f = numba.types.float3x3.as_dtype()
a = np.array([(1.0, 2.0, 3.0) for _ in range(M)], dtype=dtype_vec3f)
b = np.array([(4.0, 5.0, 6.0) for _ in range(N)], dtype=dtype_vec3f)
out = np.zeros(M * N, dtype=dtype_vec3f)

# 设备内存拷贝与内核启动
d_a = cuda.to_device(a)
d_b = cuda.to_device(b)
d_out = cuda.to_device(out)
d_SZ = cuda.to_device(SZ)

add_3darrs_broadcast[grid_size, block_size](d_out, d_a, d_b, d_SZ)

# 取回结果
res = d_out.copy_to_host()

方案2:使用(N,3)格式二维数组实现(更推荐)

该方案不需要依赖特殊向量类型,直接用标准NumPy数组存储分量,使用更符合Python习惯,性能也更稳定:

@cuda.jit
def add_3darrs_broadcast_v2(out, a, b, SZ):
    M = SZ[0]
    N = SZ[1]
    S = SZ[2]
    tx = cuda.threadIdx.x
    bx = cuda.blockIdx.x
    BSZ = cuda.blockDim.x

    for s in range(S):
        t = s * BSZ + tx
        if t < N:
            out[bx, t, 0] = b[t, 0] + a[bx, 0]
            out[bx, t, 1] = b[t, 1] + a[bx, 1]
            out[bx, t, 2] = b[t, 2] + a[bx, 2]
        cuda.syncthreads()

调用方法:

# 直接使用标准float32数组即可
a = np.full((M, 3), fill_value=[1,2,3], dtype=np.float32)
b = np.full((N, 3), fill_value=[4,5,6], dtype=np.float32)
out = np.zeros((M, N, 3), dtype=np.float32)

# 后续设备拷贝、内核启动逻辑和方案1完全一致

内容的提问来源于stack exchange,提问作者andregalera

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 06:09:04