You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Numba实现GPU并行代码无法充分调用GPU资源仅占用20~30%问题求助

问题描述

在完成大学作业过程中,需对比同逻辑GPU并行代码的Numba与CUDA-C版本性能。CUDA-C版本运行正常,Nsight检测GPU占有率达标;将代码按规则适配为Numba版本后,任务管理器显示GPU仅占用20~30%。测试矩阵乘法等标准Numba GPU代码时运行时GPU占用正常,自定义代码数据量加倍后,GPU占用率也没有对应提升。重装Anaconda、Python、Numba后问题仍未解决。

运行环境
  • 硬件:联想Ideapad GAMING 3i笔记本,搭载GTX 1650(4GB显存)
  • 软件:通过Anaconda运行Spyder 5.1.5,Python版本3.8,Numba版本0.5.4.1
问题代码(实现平方后求和的Reduce算法)
# 注:原代码缺少cuda导入,需补充 import numba.cuda as cuda
from numba.cuda.random import create_xoroshiro128p_states, xoroshiro128p_uniform_float32
import numpy as np
from numpy import float32
import random as rnd
import sys
import time

N = 1024;

n = 32;

stride = [];

for i in range(5):
    a = n//2**(i+1);   
    stride.append(a);

stride = np.array(stride,dtype = np.int32);    


n_particles = 8*1024;
n = 32;

@cuda.jit
def sphere(d_pos,cost,n,stride):
    
    index = cuda.threadIdx.y;

    i = cuda.threadIdx.x + cuda.blockDim.x * cuda.blockIdx.x;
    
    p = cuda.blockDim.y * cuda.threadIdx.x;

    sharray = cuda.shared.array(N,float32);

    if (index < n): 

        sharray[index + p] = d_pos[index + p + N* cuda.blockIdx.x];
        sharray[index + p] *= sharray[index + p];
        
        cuda.syncthreads();

        for std in range(len(stride)): 
            
                if (index < stride[std]): 
                        sharray[index + p] += sharray[index + p + stride[std]];
            
        
        
        cuda.syncthreads();

        if (index % n == 0):
            cost[i] = sharray[cuda.threadIdx.x* n];
            

d_pos = cuda.to_device(np.ones((n_particles*n),dtype = np.float32));
d_vel = cuda.to_device(np.ones((n_particles*n),dtype = np.float32));

cost = cuda.to_device(np.zeros(n_particles,dtype = np.float32));

B =  int(((n*n_particles-1)/1024 +1));

t0=time.time();

for i in range(2000):     
    
    sphere[(B,1),(32,32)](d_pos,cost,n,stride);
    cuda.synchronize();
    
print(cost.copy_to_host());
print(time.time()-t0);
根因分析
  1. Host端强制串行调度:循环内每次核函数启动后立即调用cuda.synchronize(),强制CPU等待GPU完成本次核函数计算后才下发下一个任务,导致GPU在两次核函数执行之间存在大量空闲时间,整体占用率偏低。
  2. 核函数参数与逻辑冗余:将固定长度的stride数组、固定值n作为参数传入核函数,设备端运行时计算len(stride)会引入额外开销;Reduce逻辑依赖多次块内同步,未使用warp级原语优化,计算效率偏低。
  3. 冗余代码带来额外开销:未使用的d_vel变量、重复的n变量定义、多余导入包都会增加不必要的显存申请和运行时开销,进一步降低GPU有效利用率。
解决方案
  1. 移除Host端循环内的不必要同步
    仅在所有迭代完成后、需要读取GPU侧cost数据前执行一次同步即可,修改循环部分代码如下:
for i in range(2000):     
    sphere[(B,1),(32,32)](d_pos,cost,n,stride)
# 所有核函数下发完成后再同步
cuda.synchronize()

修改后CUDA驱动会自动将2000次核函数任务排队下发,GPU无需等待CPU调度,可满负荷执行。
2. 优化核函数固定参数
将固定值n=32、stride的长度和数值直接定义为核函数内的常量,无需作为参数传递,减少设备端运行时计算开销。
3. 优化Reduce逻辑
针对32个元素的求和场景,可直接使用warp级同步原语numba.cuda.warp_reduce_sum替代循环+块同步的实现,大幅减少同步开销,提升核函数执行效率。
4. 清理冗余代码
删除未使用的d_vel变量、多余的导入包、重复的变量定义,减少不必要的显存申请和数据传输开销。

完成以上修改后,GPU占用率可恢复到和CUDA-C版本相近的水平。

内容的提问来源于stack exchange,提问作者Arthur Eckert Rüdiger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 04:24:01