Numba实现GPU并行代码无法充分调用GPU资源仅占用20~30%问题求助
问题描述
在完成大学作业过程中,需对比同逻辑GPU并行代码的Numba与CUDA-C版本性能。CUDA-C版本运行正常,Nsight检测GPU占有率达标;将代码按规则适配为Numba版本后,任务管理器显示GPU仅占用20~30%。测试矩阵乘法等标准Numba GPU代码时运行时GPU占用正常,自定义代码数据量加倍后,GPU占用率也没有对应提升。重装Anaconda、Python、Numba后问题仍未解决。
运行环境
- 硬件:联想Ideapad GAMING 3i笔记本,搭载GTX 1650(4GB显存)
- 软件:通过Anaconda运行Spyder 5.1.5,Python版本3.8,Numba版本0.5.4.1
问题代码(实现平方后求和的Reduce算法)
# 注:原代码缺少cuda导入,需补充 import numba.cuda as cuda from numba.cuda.random import create_xoroshiro128p_states, xoroshiro128p_uniform_float32 import numpy as np from numpy import float32 import random as rnd import sys import time N = 1024; n = 32; stride = []; for i in range(5): a = n//2**(i+1); stride.append(a); stride = np.array(stride,dtype = np.int32); n_particles = 8*1024; n = 32; @cuda.jit def sphere(d_pos,cost,n,stride): index = cuda.threadIdx.y; i = cuda.threadIdx.x + cuda.blockDim.x * cuda.blockIdx.x; p = cuda.blockDim.y * cuda.threadIdx.x; sharray = cuda.shared.array(N,float32); if (index < n): sharray[index + p] = d_pos[index + p + N* cuda.blockIdx.x]; sharray[index + p] *= sharray[index + p]; cuda.syncthreads(); for std in range(len(stride)): if (index < stride[std]): sharray[index + p] += sharray[index + p + stride[std]]; cuda.syncthreads(); if (index % n == 0): cost[i] = sharray[cuda.threadIdx.x* n]; d_pos = cuda.to_device(np.ones((n_particles*n),dtype = np.float32)); d_vel = cuda.to_device(np.ones((n_particles*n),dtype = np.float32)); cost = cuda.to_device(np.zeros(n_particles,dtype = np.float32)); B = int(((n*n_particles-1)/1024 +1)); t0=time.time(); for i in range(2000): sphere[(B,1),(32,32)](d_pos,cost,n,stride); cuda.synchronize(); print(cost.copy_to_host()); print(time.time()-t0);
根因分析
- Host端强制串行调度:循环内每次核函数启动后立即调用
cuda.synchronize(),强制CPU等待GPU完成本次核函数计算后才下发下一个任务,导致GPU在两次核函数执行之间存在大量空闲时间,整体占用率偏低。 - 核函数参数与逻辑冗余:将固定长度的
stride数组、固定值n作为参数传入核函数,设备端运行时计算len(stride)会引入额外开销;Reduce逻辑依赖多次块内同步,未使用warp级原语优化,计算效率偏低。 - 冗余代码带来额外开销:未使用的
d_vel变量、重复的n变量定义、多余导入包都会增加不必要的显存申请和运行时开销,进一步降低GPU有效利用率。
解决方案
- 移除Host端循环内的不必要同步
仅在所有迭代完成后、需要读取GPU侧cost数据前执行一次同步即可,修改循环部分代码如下:
for i in range(2000): sphere[(B,1),(32,32)](d_pos,cost,n,stride) # 所有核函数下发完成后再同步 cuda.synchronize()
修改后CUDA驱动会自动将2000次核函数任务排队下发,GPU无需等待CPU调度,可满负荷执行。
2. 优化核函数固定参数
将固定值n=32、stride的长度和数值直接定义为核函数内的常量,无需作为参数传递,减少设备端运行时计算开销。
3. 优化Reduce逻辑
针对32个元素的求和场景,可直接使用warp级同步原语numba.cuda.warp_reduce_sum替代循环+块同步的实现,大幅减少同步开销,提升核函数执行效率。
4. 清理冗余代码
删除未使用的d_vel变量、多余的导入包、重复的变量定义,减少不必要的显存申请和数据传输开销。
完成以上修改后,GPU占用率可恢复到和CUDA-C版本相近的水平。
内容的提问来源于stack exchange,提问作者Arthur Eckert Rüdiger
相关产品推荐
相关产品推荐

