You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用torch.cuda.Stream()控制CUDA流并行任务?测试结果不符理论求解答

问题:CUDA流并行矩阵计算结果不符合理论预期

我编写了测试代码,对比PyTorch默认CUDA流与自定义流的矩阵计算耗时,定义了三个函数分别实例化1、2、4个自定义流。理论上,这三种实现与串行执行相比应分别无加速、2倍加速、4倍加速,但实际基准测试结果完全偏离预期:

  • 在NVIDIA V100和A100 GPU上,自定义流上的计算比默认流更快;
  • 在NVIDIA V100和A100 GPU上,无论分配1、2还是4个自定义流,均未观测到可衡量的并行加速;
  • 在海光DCU K100 GPU上,自定义流上的工作负载比默认流的任务延迟更高;
  • 在海光DCU K100 GPU上,增加任意数量的自定义流都无法带来并行计算加速。

怀疑测试代码存在缺陷,有人认为是计算所用矩阵尺寸过大占满CUDA核心,但将矩阵尺寸缩小至100×100后,实验结果仍与之前类似。


测试代码

import torch
import time

def test_stream_parallel1():

    # 预热
    a = torch.randn(10000, 10000, device='cuda')
    b = torch.randn(10000, 10000, device='cuda')
    c = a @ b

    # 创建大张量模拟计算
    x1 = torch.randn(10000, 10000, device='cuda')
    x2 = torch.randn(10000, 10000, device='cuda')

    # 测试串行版本
    start = time.time()
    y1 = x1 @ x1
    y2 = x2 @ x2
    y3 = x1 @ x2
    y4 = x2 @ x1
    torch.cuda.synchronize()
    elapsed_serial = time.time() - start

    stream1 = torch.cuda.Stream()

    # 创建大张量模拟计算
    x1 = torch.randn(10000, 10000, device='cuda')
    x2 = torch.randn(10000, 10000, device='cuda')
    
    # 测试并行版本
    start = time.time()
    with torch.cuda.stream(stream1):
        y1 = x1 @ x1
        y2 = x2 @ x2
        y3 = x1 @ x2
        y4 = x2 @ x1
    torch.cuda.synchronize()
    elapsed = time.time() - start

    print(f"Serial time: {elapsed_serial:.2f}s")
    print(f"Parallel time: {elapsed:.2f}s")
    print(f"Speedup: {elapsed_serial/elapsed:.2f}x")

def test_stream_parallel2():

    # 预热
    a = torch.randn(10000, 10000, device='cuda')
    b = torch.randn(10000, 10000, device='cuda')
    c = a @ b

    # 创建大张量模拟计算
    x1 = torch.randn(10000, 10000, device='cuda')
    x2 = torch.randn(10000, 10000, device='cuda')

    # 测试串行版本
    start = time.time()
    y1 = x1 @ x1
    y2 = x2 @ x2
    y3 = x1 @ x2
    y4 = x2 @ x1
    torch.cuda.synchronize()
    elapsed_serial = time.time() - start

    stream1 = torch.cuda.Stream()
    stream2 = torch.cuda.Stream()

    # 创建大张量模拟计算
    x1 = torch.randn(10000, 10000, device='cuda')
    x2 = torch.randn(10000, 10000, device='cuda')
    
    # 测试并行版本
    start = time.time()
    with torch.cuda.stream(stream1):
        y1 = x1 @ x1
        y2 = x2 @ x2
    with torch.cuda.stream(stream2):
        y3 = x1 @ x2
        y4 = x2 @ x1
    torch.cuda.synchronize()
    elapsed = time.time() - start

    print(f"Serial time: {elapsed_serial:.2f}s")
    print(f"Parallel time: {elapsed:.2f}s")
    print(f"Speedup: {elapsed_serial/elapsed:.2f}x")


def test_stream_parallel4():

    # 预热
    a = torch.randn(10000, 10000, device='cuda')
    b = torch.randn(10000, 10000, device='cuda')
    c = a @ b

    # 创建大张量模拟计算
    x1 = torch.randn(10000, 10000, device='cuda')
    x2 = torch.randn(10000, 10000, device='cuda')

    # 测试串行版本
    start = time.time()
    y1 = x1 @ x1
    y2 = x2 @ x2
    y3 = x1 @ x2
    y4 = x2 @ x1
    torch.cuda.synchronize()
    elapsed_serial = time.time() - start

    stream1 = torch.cuda.Stream()
    stream2 = torch.cuda.Stream()
    stream3 = torch.cuda.Stream()
    stream4 = torch.cuda.Stream()

    # 创建大张量模拟计算
    x1 = torch.randn(10000, 10000, device='cuda')
    x2 = torch.randn(10000, 10000, device='cuda')
    
    # 测试并行版本
    start = time.time()
    with torch.cuda.stream(stream1):
        y1 = x1 @ x1
    with torch.cuda.stream(stream2):
        y2 = x2 @ x2
    with torch.cuda.stream(stream3):
        y3 = x1 @ x2
    with torch.cuda.stream(stream4):
        y4 = x2 @ x1
    torch.cuda.synchronize()
    elapsed = time.time() - start

    print(f"Serial time: {elapsed_serial:.2f}s")
    print(f"Parallel time: {elapsed:.2f}s")
    print(f"Speedup: {elapsed_serial/elapsed:.2f}x")



test_stream_parallel1()
print("--"*20)
test_stream_parallel2()
print("--"*20)
test_stream_parallel4()
print("--"*20)

实验结果

NVIDIA v100-16G python-3.11 Cuda-12.4.1 torch 2.5.1
----------------------------------------
Serial time: 0.76s
Parallel time: 0.60s
Speedup: 1.26x
----------------------------------------
Serial time: 0.75s
Parallel time: 0.60s
Speedup: 1.25x
----------------------------------------
Serial time: 0.76s
Parallel time: 0.60s
Speedup: 1.26x
----------------------------------------

NVIDIA A100G-40G python-3.11 Cuda-12.1.1 torch 2.4.0
----------------------------------------
Serial time: 0.69s
Parallel time: 0.47s
Speedup: 1.47x
----------------------------------------
Serial time: 0.59s
Parallel time: 0.47s
Speedup: 1.25x
----------------------------------------
Serial time: 0.59s
Parallel time: 0.47s
Speedup: 1.25x
----------------------------------------

海光 DCU K100-48G python-3.8.0 torch 2.1.0
----------------------------------------
Serial time: 0.25s
Parallel time: 0.36s
Speedup: 0.68x
----------------------------------------
Serial time: 0.31s
Parallel time: 0.72s
Speedup: 0.42x
----------------------------------------
Serial time: 0.31s
Parallel time: 0.75s
Speedup: 0.41x
----------------------------------------

内容的提问来源于stack exchange,提问作者user32801042

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.05 07:14:51