PyTorch多GPU异步计算未生效,模型串行运行问题排查
多GPU部署模型无法并行运行问题分析
我构建了LargefcNet全连接大模型,分别部署在cuda:0、cuda:1、cuda:2三个GPU设备上,并编写了三种多模型计算方式进行测试。但结果显示三种方式均为串行执行,总耗时约为单GPU运行的3倍,而非预期的接近单GPU耗时。请问为何异步计算未生效,模型无法并行运行?
模型定义
class LargefcNet(nn.Module): def __init__(self, input_size, hidden_size, output_size, dropout=0.2): super(LargefcNet, self).__init__() self.fc1 = nn.Linear(input_size, hidden_size) self.fc2 = nn.Linear(hidden_size, hidden_size) self.fc3 = nn.Linear(hidden_size, hidden_size) self.fc4 = nn.Linear(hidden_size, hidden_size) self.fc5 = nn.Linear(hidden_size, hidden_size) self.fc6 = nn.Linear(hidden_size, hidden_size) self.end = nn.Linear(hidden_size, output_size) self.dropout = nn.Dropout(dropout) self.relu = nn.ReLU() def forward(self, x): x = self.dropout(self.relu(self.fc1(x))) x = self.dropout(self.relu(self.fc2(x))) x = self.dropout(self.relu(self.fc3(x))) x = self.dropout(self.relu(self.fc4(x))) x = self.dropout(self.relu(self.fc5(x))) x = self.dropout(self.relu(self.fc6(x))) x = self.end(x) return x
设备部署与测试代码
model1 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:0')) model2 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:1')) model3 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:2')) input1 = tc.randn(100, 100).to(tc.device('cuda:0')) input2 = tc.randn(100, 100).to(tc.device('cuda:1')) input3 = tc.randn(100, 100).to(tc.device('cuda:2')) # GPU加载 output1 = model1(input1) output2 = model2(input2) output3 = model3(input3) # 单GPU测试 start_time = time.time() for i in range(10): output1 = model1(input1) print(f'output1: {time.time() - start_time}') start_time = time.time() for i in range(10): output2 = model2(input2) print(f'output2: {time.time() - start_time}') start_time = time.time() for i in range(10): output3 = model3(input3) print(f'output3: {time.time() - start_time}') # 方法1 start_time = time.time() for i in range(10): model1(input1) model2(input2) model3(input3) print(f'output1, output2, output3: {time.time() - start_time}') # 方法2 start_time = time.time() for i in range(10): model1(input1).to(tc.device('cuda:0')) model2(input2).to(tc.device('cuda:1')) model3(input3).to(tc.device('cuda:2')) print(f'output1, output2, output3 with to: {time.time() - start_time}') # 方法3 start_time = time.time() for i in range(10): outputs = [model(input) for model, input in zip([model1, model2, model3], [input1, input2, input3])] print(f'outputs: {time.time() - start_time}')
测试结果
output1: 0.13068866729736328 output2: 0.13286447525024414 output3: 0.13341188430786133 output1, output2, output3: 0.37032580375671387 output1, output2, output3 with to: 0.366225004196167 outputs: 0.36612439155578613
我预期方法1、2的耗时应在0.13~0.2秒左右,但实际接近单GPU耗时的3倍,说明模型是串行运行而非并行。
问题原因
PyTorch中CUDA操作默认是异步的,但你的代码里没有显式调度异步执行,且Python主线程会隐式等待每个CUDA任务完成:
- 调用
model(input)时,PyTorch会把任务提交到对应GPU的默认流,但Python线程在执行下一行代码前,会因为要确保返回的CUDA张量处于可用状态,隐式等待当前GPU任务完成,导致三个GPU任务被串行执行。
解决方法
要实现真正的并行,需要利用CUDA流(Stream) 来独立调度不同GPU的任务,避免主线程等待单个任务完成。修改后的测试代码如下:
import torch as tc import time import torch.nn as nn class LargefcNet(nn.Module): def __init__(self, input_size, hidden_size, output_size, dropout=0.2): super(LargefcNet, self).__init__() self.fc1 = nn.Linear(input_size, hidden_size) self.fc2 = nn.Linear(hidden_size, hidden_size) self.fc3 = nn.Linear(hidden_size, hidden_size) self.fc4 = nn.Linear(hidden_size, hidden_size) self.fc5 = nn.Linear(hidden_size, hidden_size) self.fc6 = nn.Linear(hidden_size, hidden_size) self.end = nn.Linear(hidden_size, output_size) self.dropout = nn.Dropout(dropout) self.relu = nn.ReLU() def forward(self, x): x = self.dropout(self.relu(self.fc1(x))) x = self.dropout(self.relu(self.fc2(x))) x = self.dropout(self.relu(self.fc3(x))) x = self.dropout(self.relu(self.fc4(x))) x = self.dropout(self.relu(self.fc5(x))) x = self.dropout(self.relu(self.fc6(x))) x = self.end(x) return x model1 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:0')) model2 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:1')) model3 = LargefcNet(100, 10000, 100, dropout=0.4).to(tc.device('cuda:2')) input1 = tc.randn(100, 100).to(tc.device('cuda:0')) input2 = tc.randn(100, 100).to(tc.device('cuda:1')) input3 = tc.randn(100, 100).to(tc.device('cuda:2')) # 为每个GPU创建独立的CUDA流 stream0 = tc.cuda.Stream(device='cuda:0') stream1 = tc.cuda.Stream(device='cuda:1') stream2 = tc.cuda.Stream(device='cuda:2') # 并行执行测试 start_time = time.time() for i in range(10): # 在各自流中异步提交任务 with tc.cuda.stream(stream0): output1 = model1(input1) with tc.cuda.stream(stream1): output2 = model2(input2) with tc.cuda.stream(stream2): output3 = model3(input3) # 等待所有GPU任务完成后再计时 tc.cuda.synchronize() print(f'并行执行耗时: {time.time() - start_time}')
关键说明
- CUDA流的作用:每个GPU的流是独立的任务队列,使用
with tc.cuda.stream(stream)上下文可以将任务提交到对应流,主线程无需等待单个任务完成,就能连续提交三个GPU任务,实现并行调度。 - 同步操作的必要性:
tc.cuda.synchronize()会等待所有GPU任务执行完毕后再计算总耗时,避免因为主线程先完成而导致计时不准确。 - 冗余操作移除:原代码中方法2的
.to(tc.device('cuda:x'))完全多余,因为模型和输入已经在对应GPU上,该操作会触发同步,反而增加额外耗时。
内容的提问来源于stack exchange,提问作者user15299229
相关产品推荐
相关产品推荐

