PyTorch多GPU训练报错:CUDA设备繁忙或不可用,求解决方案
问题
使用PyTorch的DistributedDataParallel实现目标检测器多GPU训练,通过torch.multiprocessing.spawn启动主进程时,调用DistributedDataParallel出现如下错误:
-- Process 1 terminated with the following error: Traceback (most recent call last): File "/home/user/.conda/envs/newenv/lib/python3.9/site-packages/torch/multiprocessing/spawn.py", line 69, in _wrap fn(i, *args) File "/home/ user/MA/code_copy/masterarbei/src/faster_rcnn/faster_rcnn_v1_parallel.py", line 178, in main_initial_train model = DDP(model, device_ids=[rank]) File "/home/user/.conda/envs/newenv/lib/python3.9/site-packages/torch/nn/parallel/distributed.py", line 674, in __init__ _verify_param_shape_across_processes(self.process_group, parameters) File "/home/user/.conda/envs/newenv/lib/python3.9/site-packages/torch/distributed/utils.py", line 118, in _verify_param_shape_across_processes return dist._verify_params_across_processes(process_group, tensors, logger) RuntimeError: CUDA error: CUDA-capable device(s) is/are busy or unavailable CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1. Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
已通过world_size = torch.cuda.device_count()设置可用GPU数量,理论上所有GPU都应可用,求问题原因及解决方法。
原因分析
- 进程-GPU绑定混乱:
spawn启动多进程时,若未明确绑定每个进程到对应GPU,多个进程会争抢同一GPU资源,导致设备繁忙。 - 模型未正确分配到目标GPU:初始化DDP前,模型没有被移动到当前进程对应的GPU上,触发跨设备参数校验失败,间接报出设备繁忙错误。
- 残留进程占用GPU:之前的训练进程未完全终止,残留进程霸占GPU资源,新进程无法获取可用设备。
- 异步错误误导:报错提示的“设备繁忙”可能是其他异步CUDA操作(比如参数初始化、数据加载)出错的间接结果,并非真的GPU硬件繁忙。
解决方法
强制绑定进程与GPU
在子进程的入口函数开头,先指定当前进程使用的GPU:torch.cuda.set_device(rank)确保每个进程只操作自己对应的GPU,避免资源抢占。
先移模型到GPU再初始化DDP
创建DDP实例前,必须把模型移动到当前rank对应的GPU上:model = model.to(rank) model = DDP(model, device_ids=[rank])清理残留GPU进程
执行命令强制杀掉所有占用GPU的Python进程:ps aux | grep python | grep -v grep | awk '{print $2}' | xargs kill -9清理完成后再重新启动训练。
启用同步调试定位真实错误
按照报错提示设置环境变量,让CUDA操作同步执行,获取准确的错误栈:CUDA_LAUNCH_BLOCKING=1 python your_train_script.py这样能排查出参数初始化、数据加载等环节的真实问题。
检查分布式进程组初始化
确保子进程中正确初始化分布式进程组,示例代码:torch.distributed.init_process_group( backend='nccl', init_method='tcp://127.0.0.1:23456', world_size=world_size, rank=rank )进程组未正确初始化会导致设备资源分配异常。
内容的提问来源于stack exchange,提问作者MS Du
相关产品推荐
相关产品推荐

