You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CUDA Python神经网络隐藏层代码优化求助:运行时长未达标

CUDA Python代码优化求助:神经网络隐藏层GPU加速未达性能要求

我正在学习英伟达《Fundamentals of Accelerated Computing with CUDA Python》课程,需完成神经网络隐藏层代码的重构任务。已使用Numba的@vectorize装饰器将相关函数迁移至CUDA GPU运行,但当前代码运行时长为1.23s,未达到任务要求的1s以内,恳请帮忙排查问题并提供优化方案。

原始代码

import numpy as np
from numba import cuda, vectorize

n = 1000000

greyscales = np.floor(np.random.uniform(0, 255, n).astype(np.float32))
weights = np.random.normal(.5, .1, n).astype(np.float32)

from numpy import exp

def normalize(grayscales):
    return grayscales / 255

def weigh(values, weights):
    return values * weights
    
def activate(values):
    return ( exp(values) - exp(-values) ) / ( exp(values) + exp(-values) )

def create_hidden_layer(n, greyscales, weights, exp, normalize, weigh, activate):
    normalized = normalize(greyscales)
    weighted = weigh(normalized, weights)
    activated = activate(weighted)
    return activated

arguments = {"n":n,
            "greyscales": greyscales,
            "weights": weights,
            "exp": exp,
            "normalize": normalize,
            "weigh": weigh,
            "activate": activate}

a = create_hidden_layer(**arguments)
print(a)

修改后的GPU版本代码

from math import exp

@vectorize(['float32(float32)'],target='cuda')
def normalize(grayscales):
    return grayscales / 255

@vectorize(['float32(float32,float32)'],target='cuda')
def weigh(values, weights):
    return values * weights

@vectorize(['float32(float32)'],target='cuda')
def activate(values):
    return ( exp(values) - exp(-values) ) / ( exp(values) + exp(-values) )

def create_hidden_layer(n, greyscales, weights, exp, normalize, weigh, activate):
    normalized = normalize(greyscales)
    weighted = weigh(normalized, weights)
    activated = activate(weighted)
    return activated

greyscales = cuda.to_device(greyscales)
weights = cuda.to_device(weights)

normalized = cuda.device_array(shape=(n,), dtype=np.float32)
weighted = cuda.device_array(shape=(n,), dtype=np.float32)
activated = cuda.device_array(shape=(n,), dtype=np.float32)

activated = activated.copy_to_host()

arguments = {"n":n,
            "greyscales": greyscales,
            "weights": weights,
            "exp": exp,
            "normalize": normalize,
            "weigh": weigh,
            "activate": activate}

a = create_hidden_layer(**arguments)
print(a)

问题排查

  • 冗余的设备内存操作:手动创建的normalized、weighted、activated设备数组并未被实际使用,@vectorize装饰的函数会自动分配内存处理数据;提前执行的activated.copy_to_host()是完全多余的操作,平白增加了数据传输时间。
  • 激活函数计算冗余:手动实现双曲正切函数时重复计算了exp(values)和exp(-values),且调用math.exp的效率不如Numba针对CUDA优化的内置数学函数。
  • 冗余参数传递:create_hidden_layer函数中的exp参数并未被GPU版本的activate函数使用,属于不必要的开销。
  • 中间内存读写开销:三次独立的vectorize函数调用会产生三次中间数组的内存读写,增加了延迟。

优化方案

1. 移除冗余操作

删除未使用的设备数组创建代码和提前的copy_to_host调用,减少无意义的内存操作和数据传输。

2. 替换激活函数为内置tanh

Numba的CUDA环境提供了优化的tanh内置函数,直接调用可以避免重复计算指数,提升计算效率:

from numba import tanh

@vectorize(['float32(float32)'],target='cuda')
def activate(values):
    return tanh(values)

3. 清理冗余参数

移除create_hidden_layer函数参数列表中的exp,以及调用时传入的对应参数,减少不必要的参数传递。

4. 合并操作减少内存读写

将normalize、weigh、activate三个步骤合并为一个CUDA kernel,避免中间数组的内存读写开销。示例如下:

from numba import cuda, tanh

@cuda.jit
def hidden_layer_kernel(greyscales, weights, output):
    idx = cuda.grid(1)
    if idx < greyscales.size:
        normalized = greyscales[idx] / 255
        weighted = normalized * weights[idx]
        output[idx] = tanh(weighted)

# 调用方式
threads_per_block = 256
blocks_per_grid = (n + threads_per_block - 1) // threads_per_block

output = cuda.device_array(n, dtype=np.float32)
hidden_layer_kernel[blocks_per_grid, threads_per_block](greyscales, weights, output)
a = output.copy_to_host()

优化后的完整代码

import numpy as np
from numba import cuda, tanh

n = 1000000

# 生成数据
greyscales = np.floor(np.random.uniform(0, 255, n).astype(np.float32))
weights = np.random.normal(.5, .1, n).astype(np.float32)

# 传输数据到GPU
greyscales_dev = cuda.to_device(greyscales)
weights_dev = cuda.to_device(weights)
output_dev = cuda.device_array(n, dtype=np.float32)

# 合并式CUDA Kernel
@cuda.jit
def hidden_layer_kernel(greyscales, weights, output):
    idx = cuda.grid(1)
    if idx < greyscales.size:
        normalized = greyscales[idx] / 255
        weighted = normalized * weights[idx]
        output[idx] = tanh(weighted)

# 配置并启动Kernel
threads_per_block = 256
blocks_per_grid = (n + threads_per_block - 1) // threads_per_block
hidden_layer_kernel[blocks_per_grid, threads_per_block](greyscales_dev, weights_dev, output_dev)

# 结果回传
a = output_dev.copy_to_host()
print(a)

内容的提问来源于stack exchange,提问作者kndrtt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 22:15:38