使用SLURM运行PyTorch分布式训练报“地址族不支持”错误求助
问题描述
在GPU集群的2个节点(每节点2张V100 GPU)上,通过SLURM脚本调用torch.distributed.run运行分布式Python代码时,出现socket初始化错误:
[W socket.cpp:426] [c10d] The server socket cannot be initialized on [::]:16773 (errno: 97 - Address family not supported by protocol). [W socket.cpp:601] [c10d] The client socket cannot be initialized to connect to [clara06.url.de]:16773 (errno: 97 - Address family not supported by protocol).
所用SLURM脚本:
#!/bin/bash #SBATCH --job-name=distribution-test # name #SBATCH --nodes=2 # nodes #SBATCH --ntasks-per-node=1 # crucial - only 1 task per dist per node! #SBATCH --cpus-per-task=4 # number of cores per tasks #SBATCH --partition=clara #SBATCH --gres=gpu:v100:2 # number of gpus #SBATCH --time 0:15:00 # maximum execution time (HH:MM:SS) #SBATCH --output=%x-%j.out # output file name module load Python pip install --user -r requirements.txt MASTER_ADDR=$(scontrol show hostnames "$SLURM_JOB_NODELIST" | head -n 1) MASTER_PORT=$(expr 10000 + $(echo -n $SLURM_JOBID | tail -c 4)) GPUS_PER_NODE=2 LOGLEVEL=INFO python -m torch.distributed.run --rdzv_id=$SLURM_JOBID --rdzv_backend=c10d --rdzv_endpoint=$MASTER_ADDR\:$MASTER_PORT --nproc_per_node $GPUS_PER_NODE --nnodes $SLURM_NNODES torch-distributed-gpu-test.py
待运行Python代码:
import fcntl import os import socket import torch import torch.distributed as dist def printflock(*msgs): """solves multi-process interleaved print problem""" with open(__file__, "r") as fh: fcntl.flock(fh, fcntl.LOCK_EX) try: print(*msgs) finally: fcntl.flock(fh, fcntl.LOCK_UN) local_rank = int(os.environ["LOCAL_RANK"]) torch.cuda.set_device(local_rank) device = torch.device("cuda", local_rank) hostname = socket.gethostname() gpu = f"[{hostname}-{local_rank}]" try: # test distributed dist.init_process_group("nccl") dist.all_reduce(torch.ones(1).to(device), op=dist.ReduceOp.SUM) dist.barrier() # test cuda is available and can allocate memory torch.cuda.is_available() torch.ones(1).cuda(local_rank) # global rank rank = dist.get_rank() world_size = dist.get_world_size() printflock(f"{gpu} is OK (global rank: {rank}/{world_size})") dist.barrier() if rank == 0: printflock(f"pt={torch.__version__}, cuda={torch.version.cuda}, nccl={torch.cuda.nccl.version()}") except Exception: printflock(f"{gpu} is broken") raise
已尝试操作:
- 切换
torch.distributed.run(指定master地址端口)、torchrun、torch.distributed.launch等启动命令 - 显式指定IP地址而非主机名
- 检查端口可用性
- 验证主机名到IP的映射
- 测试节点间连通性
- 尝试强制指定IPv4
以上操作均未解决问题。
解决方案
1. 强制禁用IPv6,指定网卡
错误核心是PyTorch默认尝试使用IPv6,但集群环境不支持。通过环境变量强制使用IPv4和指定集群可用网卡:
修改SLURM脚本中的启动命令,添加环境变量:
LOGLEVEL=info ADDR_INET6=0 NCCL_SOCKET_IFNAME=eth0 python -m torch.distributed.run --rdzv_id=$SLURM_JOBID --rdzv_backend=c10d --rdzv_endpoint=$MASTER_ADDR\:$MASTER_PORT --nproc_per_node $GPUS_PER_NODE --nnodes $SLURM_NNODES torch-distributed-gpu-test.py
ADDR_INET6=0:强制PyTorch禁用IPv6NCCL_SOCKET_IFNAME=eth0:替换为集群实际使用的网卡名称(可通过ip addr查看,比如ib0、eno1等)
2. 调整SLURM任务配置
在SBATCH参数中添加环境变量传递和内存配置,确保进程能获取完整环境:
#SBATCH --export=ALL #SBATCH --mem-per-cpu=4G # 根据实际需求调整内存大小
同时在脚本中显式导出MASTER地址和端口,确保所有进程能读取到:
export MASTER_ADDR=$MASTER_ADDR export MASTER_PORT=$MASTER_PORT
3. 显式初始化分布式进程组
修改Python代码中的dist.init_process_group,显式指定TCP初始化方式,避免自动选择IPv6:
# 替换原有的dist.init_process_group("nccl") dist.init_process_group( backend="nccl", init_method=f"tcp://{os.environ['MASTER_ADDR']}:{os.environ['MASTER_PORT']}", rank=int(os.environ['RANK']), world_size=int(os.environ['WORLD_SIZE']) )
4. 验证依赖版本兼容性
运行以下命令检查PyTorch、CUDA、NCCL版本是否匹配:
python -c "import torch; print(f'PyTorch: {torch.__version__}'); print(f'CUDA: {torch.version.cuda}'); print(f'NCCL: {torch.cuda.nccl.version()}')"
如果版本不兼容,需要升级NCCL或切换到匹配的PyTorch版本。
5. 联系集群管理员排查网络
如果以上方案都无效,联系集群管理员确认:
- 节点间防火墙是否开放了10000-20000端口范围(或你指定的MASTER_PORT范围)
- 集群是否全局禁用了IPv6
- 网卡驱动、网络服务是否正常运行
内容的提问来源于stack exchange,提问作者Scorix
相关产品推荐
相关产品推荐

