You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用SLURM运行PyTorch分布式训练报“地址族不支持”错误求助

问题描述

在GPU集群的2个节点(每节点2张V100 GPU)上,通过SLURM脚本调用torch.distributed.run运行分布式Python代码时,出现socket初始化错误:

[W socket.cpp:426] [c10d] The server socket cannot be initialized on [::]:16773 (errno: 97 - Address family not supported by protocol).
[W socket.cpp:601] [c10d] The client socket cannot be initialized to connect to [clara06.url.de]:16773 (errno: 97 - Address family not supported by protocol).

所用SLURM脚本:

#!/bin/bash
#SBATCH --job-name=distribution-test        # name
#SBATCH --nodes=2                           # nodes
#SBATCH --ntasks-per-node=1                 # crucial - only 1 task per dist per node!
#SBATCH --cpus-per-task=4                   # number of cores per tasks
#SBATCH --partition=clara
#SBATCH --gres=gpu:v100:2                   # number of gpus
#SBATCH --time 0:15:00                      # maximum execution time (HH:MM:SS)
#SBATCH --output=%x-%j.out                  # output file name

module load Python
pip install --user -r requirements.txt
MASTER_ADDR=$(scontrol show hostnames "$SLURM_JOB_NODELIST" | head -n 1)
MASTER_PORT=$(expr 10000 + $(echo -n $SLURM_JOBID | tail -c 4))
GPUS_PER_NODE=2

LOGLEVEL=INFO python -m torch.distributed.run --rdzv_id=$SLURM_JOBID --rdzv_backend=c10d --rdzv_endpoint=$MASTER_ADDR\:$MASTER_PORT --nproc_per_node $GPUS_PER_NODE --nnodes $SLURM_NNODES  torch-distributed-gpu-test.py

待运行Python代码:

import fcntl
import os
import socket

import torch
import torch.distributed as dist


def printflock(*msgs):
    """solves multi-process interleaved print problem"""
    with open(__file__, "r") as fh:
        fcntl.flock(fh, fcntl.LOCK_EX)
        try:
            print(*msgs)
        finally:
            fcntl.flock(fh, fcntl.LOCK_UN)


local_rank = int(os.environ["LOCAL_RANK"])
torch.cuda.set_device(local_rank)
device = torch.device("cuda", local_rank)
hostname = socket.gethostname()

gpu = f"[{hostname}-{local_rank}]"

try:
    # test distributed
    dist.init_process_group("nccl")
    dist.all_reduce(torch.ones(1).to(device), op=dist.ReduceOp.SUM)
    dist.barrier()

    # test cuda is available and can allocate memory
    torch.cuda.is_available()
    torch.ones(1).cuda(local_rank)

    # global rank
    rank = dist.get_rank()
    world_size = dist.get_world_size()

    printflock(f"{gpu} is OK (global rank: {rank}/{world_size})")

    dist.barrier()
    if rank == 0:
        printflock(f"pt={torch.__version__}, cuda={torch.version.cuda}, nccl={torch.cuda.nccl.version()}")

except Exception:
    printflock(f"{gpu} is broken")
    raise

已尝试操作:

  • 切换torch.distributed.run(指定master地址端口)、torchrun、torch.distributed.launch等启动命令
  • 显式指定IP地址而非主机名
  • 检查端口可用性
  • 验证主机名到IP的映射
  • 测试节点间连通性
  • 尝试强制指定IPv4

以上操作均未解决问题。

解决方案

1. 强制禁用IPv6,指定网卡

错误核心是PyTorch默认尝试使用IPv6,但集群环境不支持。通过环境变量强制使用IPv4和指定集群可用网卡:

修改SLURM脚本中的启动命令,添加环境变量:

LOGLEVEL=info ADDR_INET6=0 NCCL_SOCKET_IFNAME=eth0 python -m torch.distributed.run --rdzv_id=$SLURM_JOBID --rdzv_backend=c10d --rdzv_endpoint=$MASTER_ADDR\:$MASTER_PORT --nproc_per_node $GPUS_PER_NODE --nnodes $SLURM_NNODES torch-distributed-gpu-test.py
  • ADDR_INET6=0:强制PyTorch禁用IPv6
  • NCCL_SOCKET_IFNAME=eth0:替换为集群实际使用的网卡名称(可通过ip addr查看,比如ib0、eno1等)

2. 调整SLURM任务配置

在SBATCH参数中添加环境变量传递和内存配置,确保进程能获取完整环境:

#SBATCH --export=ALL
#SBATCH --mem-per-cpu=4G  # 根据实际需求调整内存大小

同时在脚本中显式导出MASTER地址和端口,确保所有进程能读取到:

export MASTER_ADDR=$MASTER_ADDR
export MASTER_PORT=$MASTER_PORT

3. 显式初始化分布式进程组

修改Python代码中的dist.init_process_group,显式指定TCP初始化方式,避免自动选择IPv6:

# 替换原有的dist.init_process_group("nccl")
dist.init_process_group(
    backend="nccl",
    init_method=f"tcp://{os.environ['MASTER_ADDR']}:{os.environ['MASTER_PORT']}",
    rank=int(os.environ['RANK']),
    world_size=int(os.environ['WORLD_SIZE'])
)

4. 验证依赖版本兼容性

运行以下命令检查PyTorch、CUDA、NCCL版本是否匹配:

python -c "import torch; print(f'PyTorch: {torch.__version__}'); print(f'CUDA: {torch.version.cuda}'); print(f'NCCL: {torch.cuda.nccl.version()}')"

如果版本不兼容,需要升级NCCL或切换到匹配的PyTorch版本。

5. 联系集群管理员排查网络

如果以上方案都无效,联系集群管理员确认:

  • 节点间防火墙是否开放了10000-20000端口范围(或你指定的MASTER_PORT范围)
  • 集群是否全局禁用了IPv6
  • 网卡驱动、网络服务是否正常运行

内容的提问来源于stack exchange,提问作者Scorix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 01:18:09