You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

未设置--gres=gpu:n时Slurm未限制作业可用GPU数量的问题

解决Slurm限制作业可用GPU数量的方法

1. 修正节点GPU资源配置

你的slurm.conf中节点GPU数量配置与实际硬件不符,先修改NodeName行的GPU数量:

NodeName=gpu-server35 Gres=gpu:tesla:4 CPUs=48 RealMemory=200000 Sockets=2 CoresPerSocket=12 ThreadsPerCore=2 State=UNKNOWN

修改完成后重启Slurm服务:

systemctl restart slurmctld slurmd

2. 调整资源调度策略

将slurm.conf中的资源选择参数改为基于GPU调度:

SelectTypeParameters=CR_GPU

该配置让Slurm以GPU为核心调度单位,只有作业通过--gres=gpu:n明确申请GPU资源时,才会分配对应数量的GPU。

3. 确认Cgroup设备约束生效

你的cgroup.conf已设置ConstrainDevices=yes,此配置会让Slurm通过cgroup机制限制作业仅能访问分配到的GPU设备。确保cgroup设备子系统正常挂载,且Slurm服务拥有足够权限管理。

4. 验证效果

  • 未指定GPU的情况:运行原命令
    srun --mpi=none -n1 -p debug python test_gpu.py
    
    此时torch.cuda.device_count()应返回0,作业无法访问任何GPU。
  • 指定GPU数量的情况:运行带--gres参数的命令
    srun --mpi=none -n1 -p debug --gres=gpu:2 python test_gpu.py
    
    此时torch.cuda.device_count()应返回2,且Slurm会自动设置CUDA_VISIBLE_DEVICES环境变量为对应GPU编号。

可选:设置分区默认GPU数量

若需要默认给分区内作业分配固定数量GPU,可在PartitionName行添加DefaultGres参数:

PartitionName=debug Nodes=gpu-server35 Default=YES MaxTime=INFINITE State=UP DefaultGres=gpu:1

这样未指定--gres时,作业默认获得1块GPU。

内容的提问来源于stack exchange,提问作者kidy zhang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 16:54:57