CUDA Python神经网络隐藏层代码优化求助:运行时长未达标
CUDA Python代码优化求助:神经网络隐藏层GPU加速未达性能要求
我正在学习英伟达《Fundamentals of Accelerated Computing with CUDA Python》课程,需完成神经网络隐藏层代码的重构任务。已使用Numba的@vectorize装饰器将相关函数迁移至CUDA GPU运行,但当前代码运行时长为1.23s,未达到任务要求的1s以内,恳请帮忙排查问题并提供优化方案。
原始代码
import numpy as np from numba import cuda, vectorize n = 1000000 greyscales = np.floor(np.random.uniform(0, 255, n).astype(np.float32)) weights = np.random.normal(.5, .1, n).astype(np.float32) from numpy import exp def normalize(grayscales): return grayscales / 255 def weigh(values, weights): return values * weights def activate(values): return ( exp(values) - exp(-values) ) / ( exp(values) + exp(-values) ) def create_hidden_layer(n, greyscales, weights, exp, normalize, weigh, activate): normalized = normalize(greyscales) weighted = weigh(normalized, weights) activated = activate(weighted) return activated arguments = {"n":n, "greyscales": greyscales, "weights": weights, "exp": exp, "normalize": normalize, "weigh": weigh, "activate": activate} a = create_hidden_layer(**arguments) print(a)
修改后的GPU版本代码
from math import exp @vectorize(['float32(float32)'],target='cuda') def normalize(grayscales): return grayscales / 255 @vectorize(['float32(float32,float32)'],target='cuda') def weigh(values, weights): return values * weights @vectorize(['float32(float32)'],target='cuda') def activate(values): return ( exp(values) - exp(-values) ) / ( exp(values) + exp(-values) ) def create_hidden_layer(n, greyscales, weights, exp, normalize, weigh, activate): normalized = normalize(greyscales) weighted = weigh(normalized, weights) activated = activate(weighted) return activated greyscales = cuda.to_device(greyscales) weights = cuda.to_device(weights) normalized = cuda.device_array(shape=(n,), dtype=np.float32) weighted = cuda.device_array(shape=(n,), dtype=np.float32) activated = cuda.device_array(shape=(n,), dtype=np.float32) activated = activated.copy_to_host() arguments = {"n":n, "greyscales": greyscales, "weights": weights, "exp": exp, "normalize": normalize, "weigh": weigh, "activate": activate} a = create_hidden_layer(**arguments) print(a)
问题排查
- 冗余的设备内存操作:手动创建的
normalized、weighted、activated设备数组并未被实际使用,@vectorize装饰的函数会自动分配内存处理数据;提前执行的activated.copy_to_host()是完全多余的操作,平白增加了数据传输时间。 - 激活函数计算冗余:手动实现双曲正切函数时重复计算了
exp(values)和exp(-values),且调用math.exp的效率不如Numba针对CUDA优化的内置数学函数。 - 冗余参数传递:
create_hidden_layer函数中的exp参数并未被GPU版本的activate函数使用,属于不必要的开销。 - 中间内存读写开销:三次独立的
vectorize函数调用会产生三次中间数组的内存读写,增加了延迟。
优化方案
1. 移除冗余操作
删除未使用的设备数组创建代码和提前的copy_to_host调用,减少无意义的内存操作和数据传输。
2. 替换激活函数为内置tanh
Numba的CUDA环境提供了优化的tanh内置函数,直接调用可以避免重复计算指数,提升计算效率:
from numba import tanh @vectorize(['float32(float32)'],target='cuda') def activate(values): return tanh(values)
3. 清理冗余参数
移除create_hidden_layer函数参数列表中的exp,以及调用时传入的对应参数,减少不必要的参数传递。
4. 合并操作减少内存读写
将normalize、weigh、activate三个步骤合并为一个CUDA kernel,避免中间数组的内存读写开销。示例如下:
from numba import cuda, tanh @cuda.jit def hidden_layer_kernel(greyscales, weights, output): idx = cuda.grid(1) if idx < greyscales.size: normalized = greyscales[idx] / 255 weighted = normalized * weights[idx] output[idx] = tanh(weighted) # 调用方式 threads_per_block = 256 blocks_per_grid = (n + threads_per_block - 1) // threads_per_block output = cuda.device_array(n, dtype=np.float32) hidden_layer_kernel[blocks_per_grid, threads_per_block](greyscales, weights, output) a = output.copy_to_host()
优化后的完整代码
import numpy as np from numba import cuda, tanh n = 1000000 # 生成数据 greyscales = np.floor(np.random.uniform(0, 255, n).astype(np.float32)) weights = np.random.normal(.5, .1, n).astype(np.float32) # 传输数据到GPU greyscales_dev = cuda.to_device(greyscales) weights_dev = cuda.to_device(weights) output_dev = cuda.device_array(n, dtype=np.float32) # 合并式CUDA Kernel @cuda.jit def hidden_layer_kernel(greyscales, weights, output): idx = cuda.grid(1) if idx < greyscales.size: normalized = greyscales[idx] / 255 weighted = normalized * weights[idx] output[idx] = tanh(weighted) # 配置并启动Kernel threads_per_block = 256 blocks_per_grid = (n + threads_per_block - 1) // threads_per_block hidden_layer_kernel[blocks_per_grid, threads_per_block](greyscales_dev, weights_dev, output_dev) # 结果回传 a = output_dev.copy_to_host() print(a)
内容的提问来源于stack exchange,提问作者kndrtt
相关产品推荐
相关产品推荐

