You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch多GPU异步计算未生效,模型串行运行问题排查

多GPU部署模型无法并行运行问题分析

我构建了LargefcNet全连接大模型,分别部署在cuda:0、cuda:1、cuda:2三个GPU设备上,并编写了三种多模型计算方式进行测试。但结果显示三种方式均为串行执行,总耗时约为单GPU运行的3倍,而非预期的接近单GPU耗时。请问为何异步计算未生效,模型无法并行运行?

模型定义

class LargefcNet(nn.Module):
    def __init__(self, input_size, hidden_size, output_size, dropout=0.2):
        super(LargefcNet, self).__init__()
        self.fc1 = nn.Linear(input_size, hidden_size)
        self.fc2 = nn.Linear(hidden_size, hidden_size)
        self.fc3 = nn.Linear(hidden_size, hidden_size)
        self.fc4 = nn.Linear(hidden_size, hidden_size)
        self.fc5 = nn.Linear(hidden_size, hidden_size)
        self.fc6 = nn.Linear(hidden_size, hidden_size)
        self.end = nn.Linear(hidden_size, output_size)
        self.dropout = nn.Dropout(dropout)
        self.relu = nn.ReLU()
    def forward(self, x):
        x = self.dropout(self.relu(self.fc1(x)))
        x = self.dropout(self.relu(self.fc2(x)))
        x = self.dropout(self.relu(self.fc3(x)))
        x = self.dropout(self.relu(self.fc4(x)))
        x = self.dropout(self.relu(self.fc5(x)))
        x = self.dropout(self.relu(self.fc6(x)))
        x = self.end(x)
        return x

设备部署与测试代码

model1 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:0'))
model2 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:1'))
model3 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:2'))
input1 = tc.randn(100, 100).to(tc.device('cuda:0'))
input2 = tc.randn(100, 100).to(tc.device('cuda:1'))
input3 = tc.randn(100, 100).to(tc.device('cuda:2'))

# GPU加载
output1 = model1(input1)
output2 = model2(input2)
output3 = model3(input3)

# 单GPU测试
start_time = time.time()
for i in range(10):
    output1 = model1(input1)
print(f'output1: {time.time() - start_time}')

start_time = time.time()
for i in range(10):
    output2 = model2(input2)
print(f'output2: {time.time() - start_time}')

start_time = time.time()
for i in range(10):
    output3 = model3(input3)
print(f'output3: {time.time() - start_time}')

# 方法1
start_time = time.time()
for i in range(10):
    model1(input1)
    model2(input2)
    model3(input3)
print(f'output1, output2, output3: {time.time() - start_time}')

# 方法2
start_time = time.time()
for i in range(10):
    model1(input1).to(tc.device('cuda:0'))
    model2(input2).to(tc.device('cuda:1'))
    model3(input3).to(tc.device('cuda:2'))
print(f'output1, output2, output3 with to: {time.time() - start_time}')

# 方法3
start_time = time.time()
for i in range(10):
    outputs = [model(input) for model, input in zip([model1, model2, model3], [input1, input2, input3])]
print(f'outputs: {time.time() - start_time}')

测试结果

output1: 0.13068866729736328
output2: 0.13286447525024414
output3: 0.13341188430786133
output1, output2, output3: 0.37032580375671387
output1, output2, output3 with to: 0.366225004196167
outputs: 0.36612439155578613

我预期方法1、2的耗时应在0.13~0.2秒左右,但实际接近单GPU耗时的3倍,说明模型是串行运行而非并行。


问题原因

PyTorch中CUDA操作默认是异步的,但你的代码里没有显式调度异步执行,且Python主线程会隐式等待每个CUDA任务完成:

  • 调用model(input)时,PyTorch会把任务提交到对应GPU的默认流,但Python线程在执行下一行代码前,会因为要确保返回的CUDA张量处于可用状态,隐式等待当前GPU任务完成,导致三个GPU任务被串行执行。

解决方法

要实现真正的并行,需要利用CUDA流(Stream) 来独立调度不同GPU的任务,避免主线程等待单个任务完成。修改后的测试代码如下:

import torch as tc
import time
import torch.nn as nn

class LargefcNet(nn.Module):
    def __init__(self, input_size, hidden_size, output_size, dropout=0.2):
        super(LargefcNet, self).__init__()
        self.fc1 = nn.Linear(input_size, hidden_size)
        self.fc2 = nn.Linear(hidden_size, hidden_size)
        self.fc3 = nn.Linear(hidden_size, hidden_size)
        self.fc4 = nn.Linear(hidden_size, hidden_size)
        self.fc5 = nn.Linear(hidden_size, hidden_size)
        self.fc6 = nn.Linear(hidden_size, hidden_size)
        self.end = nn.Linear(hidden_size, output_size)
        self.dropout = nn.Dropout(dropout)
        self.relu = nn.ReLU()
    def forward(self, x):
        x = self.dropout(self.relu(self.fc1(x)))
        x = self.dropout(self.relu(self.fc2(x)))
        x = self.dropout(self.relu(self.fc3(x)))
        x = self.dropout(self.relu(self.fc4(x)))
        x = self.dropout(self.relu(self.fc5(x)))
        x = self.dropout(self.relu(self.fc6(x)))
        x = self.end(x)
        return x

model1 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:0'))
model2 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:1'))
model3 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:2'))
input1 = tc.randn(100, 100).to(tc.device('cuda:0'))
input2 = tc.randn(100, 100).to(tc.device('cuda:1'))
input3 = tc.randn(100, 100).to(tc.device('cuda:2'))

# 为每个GPU创建独立的CUDA流
stream0 = tc.cuda.Stream(device='cuda:0')
stream1 = tc.cuda.Stream(device='cuda:1')
stream2 = tc.cuda.Stream(device='cuda:2')

# 并行执行测试
start_time = time.time()
for i in range(10):
    # 在各自流中异步提交任务
    with tc.cuda.stream(stream0):
        output1 = model1(input1)
    with tc.cuda.stream(stream1):
        output2 = model2(input2)
    with tc.cuda.stream(stream2):
        output3 = model3(input3)
# 等待所有GPU任务完成后再计时
tc.cuda.synchronize()
print(f'并行执行耗时: {time.time() - start_time}')

关键说明

  • CUDA流的作用:每个GPU的流是独立的任务队列,使用with tc.cuda.stream(stream)上下文可以将任务提交到对应流,主线程无需等待单个任务完成,就能连续提交三个GPU任务,实现并行调度。
  • 同步操作的必要性:tc.cuda.synchronize()会等待所有GPU任务执行完毕后再计算总耗时,避免因为主线程先完成而导致计时不准确。
  • 冗余操作移除:原代码中方法2的.to(tc.device('cuda:x'))完全多余,因为模型和输入已经在对应GPU上,该操作会触发同步,反而增加额外耗时。

内容的提问来源于stack exchange,提问作者user15299229

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 05:16:24