Numba实现Python GPU并行计算的性能疑问:GPU为何优于受GIL限制的CPU?
关于Numba+CUDA并行计算与GPU性能优势的疑问
作为CUDA新手,我学习《Running Python script on GPU》一文后,在Colab笔记本中运行了测试代码,结果显示:无GPU时耗时3.525673059999974,使用NVIDIA Tesla T4 GPU时耗时0.07701390800002628,GPU性能远优于Intel(R) Xeon(R) CPU @ 2.20GHz。该任务并非I/O密集型,但Python的GIL(全局解释器锁)导致程序实际为单线程,因此我有两个疑问:
- 我们真的可以使用Numba库在Python中实现并行计算吗?
- 为何GPU性能优于CPU?此处的多线程是如何实现的?
测试代码
from numba import jit, cuda import numpy as np # to measure exec time from timeit import default_timer as timer # normal function to run on cpu def func(a): for i in range(10000000): a[i]+= 1 # function optimized to run on gpu @jit(target_backend='cuda') def func2(a): for i in range(10000000): a[i]+= 1 if __name__=="__main__": n = 10000000 a = np.ones(n, dtype = np.float64) start = timer() func(a) print("without GPU:", timer()-start) start = timer() func2(a) print("with GPU:", timer()-start)
系统规格
GPU: NVIDIA Tesla T4
+-----------------------------------------------------------------------------+ | NVIDIA-SMI 460.32.03 Driver Version: 460.32.03 CUDA Version: 11.2 | |-------------------------------+----------------------+----------------------+ | GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |===============================+======================+======================| | 0 Tesla T4 Off | 00000000:00:04.0 Off | 0 | | N/A 40C P8 9W / 70W | 0MiB / 15109MiB | 0% Default | | | | N/A | +-------------------------------+----------------------+----------------------+ +-----------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=============================================================================| | No running processes found | +-----------------------------------------------------------------------------+
CPU: Intel(R) Xeon(R) CPU @ 2.20GHz
processor : 0 vendor_id : GenuineIntel cpu family : 6 model : 79 model name : Intel(R) Xeon(R) CPU @ 2.20GHz stepping : 0 microcode : 0x1 cpu MHz : 2199.998 cache size : 56320 KB physical id : 0 siblings : 2 core id : 0 cpu cores : 1 apicid : 0 initial apicid : 0 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc rep_good nopl xtopology nonstop_tsc cpuid tsc_known_freq pni pclmulqdq ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch invpcid_single ssbd ibrs ibpb stibp fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm rdseed adx smap xsaveopt arat md_clear arch_capabilities bugs : cpu_meltdown spectre_v1 spectre_v2 spec_store_bypass l1tf mds swapgs taa mmio_stale_data retbleed bogomips : 4399.99 clflush size : 64 cache_alignment : 64 address sizes : 46 bits physical, 48 bits virtual power management: processor : 1 vendor_id : GenuineIntel cpu family : 6 model : 79 model name : Intel(R) Xeon(R) CPU @ 2.20GHz stepping : 0 microcode : 0x1 cpu MHz : 2199.998 cache size : 56320 KB physical id : 0 siblings : 2 core id : 0 cpu cores : 1 apicid : 1 initial apicid : 1 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc rep_good nopl xtopology nonstop_tsc cpuid tsc_known_freq pni pclmulqdq ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch invpcid_single ssbd ibrs ibpb stibp fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm rdseed adx smap xsaveopt arat md_clear arch_capabilities bugs : cpu_meltdown spectre_v1 spectre_v2 spec_store_bypass l1tf mds swapgs taa mmio_stale_data retbleed bogomips : 4399.99 clflush size : 64 cache_alignment : 64 address sizes : 46 bits physical, 48 bits virtual power management:
解答
1. 能否用Numba在Python中实现并行计算?
完全可以。Numba的核心能力就是将Python代码编译为机器码,绕开GIL的限制,支持CPU和GPU两种并行方式:
- CPU并行:使用
@njit(parallel=True)装饰器,Numba会自动识别循环中的可并行部分,利用CPU多核心执行,代码运行在GIL之外,不受单线程限制。 - GPU并行:你代码中使用的
@jit(target_backend='cuda')(更推荐用@cuda.jit)会将Python代码编译为CUDA核函数,在GPU的大量流处理器上并行执行,完全脱离Python解释器的GIL约束。
你测试中的func2就是典型的GPU并行实现,Numba会自动把循环拆解为大量并行任务,分配给GPU的多个流处理器同时执行。
2. GPU性能优于CPU的原因及多线程实现方式
性能优势的核心原因
GPU和CPU的设计目标完全不同:
- CPU是通用计算核心,追求低延迟、单线程性能,通常只有几个到几十个核心(你的Xeon仅1个物理核心、2个超线程),适合复杂逻辑、分支多的任务。
- GPU是吞吐量优化核心,追求高并行度,Tesla T4拥有2560个CUDA核心,可以同时执行数千个线程,适合你测试中这种无分支、数据独立的简单循环任务——每个数组元素的
a[i]+=1操作彼此独立,完全可以并行执行。
你的测试任务正好是GPU擅长的**单指令多数据(SIMD)**场景,GPU能一次性处理成千上万的元素,而CPU只能单线程(或双超线程)逐个处理,性能差距自然巨大。
此处的多线程实现方式
你代码中@jit(target_backend='cuda')修饰的func2被Numba编译为CUDA核函数,并行逻辑由以下步骤实现:
- Numba自动将循环拆解为线程块(Block)和线程(Thread),每个线程负责处理数组中的一个或多个元素。
- GPU调度器将这些线程分配到不同的流处理器上并行执行,所有线程同时执行同一条指令(即
a[i]+=1),这就是CUDA的**单指令多线程(SIMT)**模型。 - 整个计算过程完全在GPU硬件上执行,与Python的GIL无关——Python仅负责触发GPU任务、等待结果返回,中间计算绕开了Python解释器。
另外,更规范的GPU并行写法应该显式指定线程和块的大小,能更精准控制并行粒度,性能通常更优:
@cuda.jit def func2(a): idx = cuda.grid(1) if idx < len(a): a[idx] += 1 # 调用时指定线程块大小和网格大小 threads_per_block = 256 blocks_per_grid = (len(a) + threads_per_block - 1) // threads_per_block func2[blocks_per_grid, threads_per_block](a)
内容的提问来源于stack exchange,提问作者Sharvani
相关产品推荐
相关产品推荐

