如何利用torch.cuda.Stream()控制CUDA流并行任务?测试结果不符理论求解答
问题:CUDA流并行矩阵计算结果不符合理论预期
我编写了测试代码,对比PyTorch默认CUDA流与自定义流的矩阵计算耗时,定义了三个函数分别实例化1、2、4个自定义流。理论上,这三种实现与串行执行相比应分别无加速、2倍加速、4倍加速,但实际基准测试结果完全偏离预期:
- 在NVIDIA V100和A100 GPU上,自定义流上的计算比默认流更快;
- 在NVIDIA V100和A100 GPU上,无论分配1、2还是4个自定义流,均未观测到可衡量的并行加速;
- 在海光DCU K100 GPU上,自定义流上的工作负载比默认流的任务延迟更高;
- 在海光DCU K100 GPU上,增加任意数量的自定义流都无法带来并行计算加速。
怀疑测试代码存在缺陷,有人认为是计算所用矩阵尺寸过大占满CUDA核心,但将矩阵尺寸缩小至100×100后,实验结果仍与之前类似。
测试代码
import torch import time def test_stream_parallel1(): # 预热 a = torch.randn(10000, 10000, device='cuda') b = torch.randn(10000, 10000, device='cuda') c = a @ b # 创建大张量模拟计算 x1 = torch.randn(10000, 10000, device='cuda') x2 = torch.randn(10000, 10000, device='cuda') # 测试串行版本 start = time.time() y1 = x1 @ x1 y2 = x2 @ x2 y3 = x1 @ x2 y4 = x2 @ x1 torch.cuda.synchronize() elapsed_serial = time.time() - start stream1 = torch.cuda.Stream() # 创建大张量模拟计算 x1 = torch.randn(10000, 10000, device='cuda') x2 = torch.randn(10000, 10000, device='cuda') # 测试并行版本 start = time.time() with torch.cuda.stream(stream1): y1 = x1 @ x1 y2 = x2 @ x2 y3 = x1 @ x2 y4 = x2 @ x1 torch.cuda.synchronize() elapsed = time.time() - start print(f"Serial time: {elapsed_serial:.2f}s") print(f"Parallel time: {elapsed:.2f}s") print(f"Speedup: {elapsed_serial/elapsed:.2f}x") def test_stream_parallel2(): # 预热 a = torch.randn(10000, 10000, device='cuda') b = torch.randn(10000, 10000, device='cuda') c = a @ b # 创建大张量模拟计算 x1 = torch.randn(10000, 10000, device='cuda') x2 = torch.randn(10000, 10000, device='cuda') # 测试串行版本 start = time.time() y1 = x1 @ x1 y2 = x2 @ x2 y3 = x1 @ x2 y4 = x2 @ x1 torch.cuda.synchronize() elapsed_serial = time.time() - start stream1 = torch.cuda.Stream() stream2 = torch.cuda.Stream() # 创建大张量模拟计算 x1 = torch.randn(10000, 10000, device='cuda') x2 = torch.randn(10000, 10000, device='cuda') # 测试并行版本 start = time.time() with torch.cuda.stream(stream1): y1 = x1 @ x1 y2 = x2 @ x2 with torch.cuda.stream(stream2): y3 = x1 @ x2 y4 = x2 @ x1 torch.cuda.synchronize() elapsed = time.time() - start print(f"Serial time: {elapsed_serial:.2f}s") print(f"Parallel time: {elapsed:.2f}s") print(f"Speedup: {elapsed_serial/elapsed:.2f}x") def test_stream_parallel4(): # 预热 a = torch.randn(10000, 10000, device='cuda') b = torch.randn(10000, 10000, device='cuda') c = a @ b # 创建大张量模拟计算 x1 = torch.randn(10000, 10000, device='cuda') x2 = torch.randn(10000, 10000, device='cuda') # 测试串行版本 start = time.time() y1 = x1 @ x1 y2 = x2 @ x2 y3 = x1 @ x2 y4 = x2 @ x1 torch.cuda.synchronize() elapsed_serial = time.time() - start stream1 = torch.cuda.Stream() stream2 = torch.cuda.Stream() stream3 = torch.cuda.Stream() stream4 = torch.cuda.Stream() # 创建大张量模拟计算 x1 = torch.randn(10000, 10000, device='cuda') x2 = torch.randn(10000, 10000, device='cuda') # 测试并行版本 start = time.time() with torch.cuda.stream(stream1): y1 = x1 @ x1 with torch.cuda.stream(stream2): y2 = x2 @ x2 with torch.cuda.stream(stream3): y3 = x1 @ x2 with torch.cuda.stream(stream4): y4 = x2 @ x1 torch.cuda.synchronize() elapsed = time.time() - start print(f"Serial time: {elapsed_serial:.2f}s") print(f"Parallel time: {elapsed:.2f}s") print(f"Speedup: {elapsed_serial/elapsed:.2f}x") test_stream_parallel1() print("--"*20) test_stream_parallel2() print("--"*20) test_stream_parallel4() print("--"*20)
实验结果
NVIDIA v100-16G python-3.11 Cuda-12.4.1 torch 2.5.1 ---------------------------------------- Serial time: 0.76s Parallel time: 0.60s Speedup: 1.26x ---------------------------------------- Serial time: 0.75s Parallel time: 0.60s Speedup: 1.25x ---------------------------------------- Serial time: 0.76s Parallel time: 0.60s Speedup: 1.26x ---------------------------------------- NVIDIA A100G-40G python-3.11 Cuda-12.1.1 torch 2.4.0 ---------------------------------------- Serial time: 0.69s Parallel time: 0.47s Speedup: 1.47x ---------------------------------------- Serial time: 0.59s Parallel time: 0.47s Speedup: 1.25x ---------------------------------------- Serial time: 0.59s Parallel time: 0.47s Speedup: 1.25x ---------------------------------------- 海光 DCU K100-48G python-3.8.0 torch 2.1.0 ---------------------------------------- Serial time: 0.25s Parallel time: 0.36s Speedup: 0.68x ---------------------------------------- Serial time: 0.31s Parallel time: 0.72s Speedup: 0.42x ---------------------------------------- Serial time: 0.31s Parallel time: 0.75s Speedup: 0.41x ----------------------------------------
内容的提问来源于stack exchange,提问作者user32801042
相关产品推荐
相关产品推荐

