You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Numba实现Python GPU并行计算的性能疑问:GPU为何优于受GIL限制的CPU?

关于Numba+CUDA并行计算与GPU性能优势的疑问

作为CUDA新手,我学习《Running Python script on GPU》一文后,在Colab笔记本中运行了测试代码,结果显示:无GPU时耗时3.525673059999974,使用NVIDIA Tesla T4 GPU时耗时0.07701390800002628,GPU性能远优于Intel(R) Xeon(R) CPU @ 2.20GHz。该任务并非I/O密集型,但Python的GIL(全局解释器锁)导致程序实际为单线程,因此我有两个疑问:

  1. 我们真的可以使用Numba库在Python中实现并行计算吗?
  2. 为何GPU性能优于CPU?此处的多线程是如何实现的?

测试代码

from numba import jit, cuda
import numpy as np

# to measure exec time
from timeit import default_timer as timer

# normal function to run on cpu
def func(a):                                
    for i in range(10000000):
        a[i]+= 1    

# function optimized to run on gpu
@jit(target_backend='cuda')                     
def func2(a):
    for i in range(10000000):
        a[i]+= 1
if __name__=="__main__":
    n = 10000000                            
    a = np.ones(n, dtype = np.float64)
    
    start = timer()
    func(a)
    print("without GPU:", timer()-start)    
    
    start = timer()
    func2(a)
    print("with GPU:", timer()-start)

系统规格

GPU: NVIDIA Tesla T4

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 460.32.03    Driver Version: 460.32.03    CUDA Version: 11.2     |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|                               |                      |               MIG M. |
|===============================+======================+======================|
|   0  Tesla T4            Off  | 00000000:00:04.0 Off |                    0 |
| N/A   40C    P8     9W /  70W |      0MiB / 15109MiB |      0%      Default |
|                               |                      |                  N/A |
+-------------------------------+----------------------+----------------------+
                                                                               
+-----------------------------------------------------------------------------+
| Processes:                                                                  |
|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |
|        ID   ID                                                   Usage      |
|=============================================================================|
|  No running processes found                                                 |
+-----------------------------------------------------------------------------+

CPU: Intel(R) Xeon(R) CPU @ 2.20GHz

processor   : 0
vendor_id   : GenuineIntel
cpu family  : 6
model       : 79
model name  : Intel(R) Xeon(R) CPU @ 2.20GHz
stepping    : 0
microcode   : 0x1
cpu MHz     : 2199.998
cache size  : 56320 KB
physical id : 0
siblings    : 2
core id     : 0
cpu cores   : 1
apicid      : 0
initial apicid  : 0
fpu     : yes
fpu_exception   : yes
cpuid level : 13
wp      : yes
flags       : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc rep_good nopl xtopology nonstop_tsc cpuid tsc_known_freq pni pclmulqdq ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch invpcid_single ssbd ibrs ibpb stibp fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm rdseed adx smap xsaveopt arat md_clear arch_capabilities
bugs        : cpu_meltdown spectre_v1 spectre_v2 spec_store_bypass l1tf mds swapgs taa mmio_stale_data retbleed
bogomips    : 4399.99
clflush size    : 64
cache_alignment : 64
address sizes   : 46 bits physical, 48 bits virtual
power management:

processor   : 1
vendor_id   : GenuineIntel
cpu family  : 6
model       : 79
model name  : Intel(R) Xeon(R) CPU @ 2.20GHz
stepping    : 0
microcode   : 0x1
cpu MHz     : 2199.998
cache size  : 56320 KB
physical id : 0
siblings    : 2
core id     : 0
cpu cores   : 1
apicid      : 1
initial apicid  : 1
fpu     : yes
fpu_exception   : yes
cpuid level : 13
wp      : yes
flags       : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc rep_good nopl xtopology nonstop_tsc cpuid tsc_known_freq pni pclmulqdq ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch invpcid_single ssbd ibrs ibpb stibp fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm rdseed adx smap xsaveopt arat md_clear arch_capabilities
bugs        : cpu_meltdown spectre_v1 spectre_v2 spec_store_bypass l1tf mds swapgs taa mmio_stale_data retbleed
bogomips    : 4399.99
clflush size    : 64
cache_alignment : 64
address sizes   : 46 bits physical, 48 bits virtual
power management:

解答

1. 能否用Numba在Python中实现并行计算?

完全可以。Numba的核心能力就是将Python代码编译为机器码,绕开GIL的限制,支持CPU和GPU两种并行方式:

  • CPU并行:使用@njit(parallel=True)装饰器,Numba会自动识别循环中的可并行部分,利用CPU多核心执行,代码运行在GIL之外,不受单线程限制。
  • GPU并行:你代码中使用的@jit(target_backend='cuda')(更推荐用@cuda.jit)会将Python代码编译为CUDA核函数,在GPU的大量流处理器上并行执行,完全脱离Python解释器的GIL约束。

你测试中的func2就是典型的GPU并行实现,Numba会自动把循环拆解为大量并行任务,分配给GPU的多个流处理器同时执行。

2. GPU性能优于CPU的原因及多线程实现方式

性能优势的核心原因

GPU和CPU的设计目标完全不同:

  • CPU是通用计算核心,追求低延迟、单线程性能,通常只有几个到几十个核心(你的Xeon仅1个物理核心、2个超线程),适合复杂逻辑、分支多的任务。
  • GPU是吞吐量优化核心,追求高并行度,Tesla T4拥有2560个CUDA核心,可以同时执行数千个线程,适合你测试中这种无分支、数据独立的简单循环任务——每个数组元素的a[i]+=1操作彼此独立,完全可以并行执行。

你的测试任务正好是GPU擅长的**单指令多数据(SIMD)**场景,GPU能一次性处理成千上万的元素,而CPU只能单线程(或双超线程)逐个处理,性能差距自然巨大。

此处的多线程实现方式

你代码中@jit(target_backend='cuda')修饰的func2被Numba编译为CUDA核函数,并行逻辑由以下步骤实现:

  1. Numba自动将循环拆解为线程块(Block)和线程(Thread),每个线程负责处理数组中的一个或多个元素。
  2. GPU调度器将这些线程分配到不同的流处理器上并行执行,所有线程同时执行同一条指令(即a[i]+=1),这就是CUDA的**单指令多线程(SIMT)**模型。
  3. 整个计算过程完全在GPU硬件上执行,与Python的GIL无关——Python仅负责触发GPU任务、等待结果返回,中间计算绕开了Python解释器。

另外,更规范的GPU并行写法应该显式指定线程和块的大小,能更精准控制并行粒度,性能通常更优:

@cuda.jit
def func2(a):
    idx = cuda.grid(1)
    if idx < len(a):
        a[idx] += 1

# 调用时指定线程块大小和网格大小
threads_per_block = 256
blocks_per_grid = (len(a) + threads_per_block - 1) // threads_per_block
func2[blocks_per_grid, threads_per_block](a)

内容的提问来源于stack exchange,提问作者Sharvani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 18:20:38