You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Slurm集群中Ray分布式HPO的GPU检测异常问题求助

问题

在Slurm集群上进行分布式HPO(超参数优化)时,Ray无法正确检测GPU资源。集群包含仅运行调度器的CPU头节点,以及多台配置相同、各带4块GPU的工作节点,但Ray仅能识别其中一个节点的全部4块GPU,其余节点仅识别1块GPU。

集群启动配置

  • 头节点启动命令:
ray start --head --temp-dir /tmp/UID/ray/ --node-ip-address=10.13.22.34 --port=6374 --block --dashboard-port 3232 --dashboard-host 0.0.0.0 --num-cpus 2 &
  • 工作节点启动脚本(sbatch):
srun --gres=gpu:4 --nodes=1 --ntasks=1 -w "\${node_i}" ray start --address 10.13.22.34:6374 --block &

Ray资源检测输出

ray.nodes()和ray.cluster_resources()的输出:

[{'NodeID': '3b793f590f23b9a3f8eb4da86f870a2c6a20967ea1ea1093945fb2d9', 'Alive': True, 'NodeManagerAddress': '10.13.31.131', 'NodeManagerHostname': 'jwb0801.juwels', 'NodeManagerPort': 44655, 'ObjectManagerPort': 40891, 'ObjectStoreSocketName': '/tmp/ray/session_2023-04-05_09-43-27_453188_1518/sockets/plasma_store', 'RayletSocketName': '/tmp/ray/session_2023-04-05_09-43-27_453188_1518/sockets/raylet', 'MetricsExportPort': 43462, 'NodeName': '10.13.31.131', 'alive': True, 'Resources': {'CPU': 96.0, 'memory': 359211870618.0, 'GPU': 4.0, 'node:10.13.31.131': 1.0, 'object_store_memory': 153947944550.0, 'accelerator_type:A100': 1.0}}, {'NodeID': '5f803a7dd410074ea9d3fb3058b3ab3044011c54536c7808bd9a947c', 'Alive': True, 'NodeManagerAddress': '10.13.22.34', 'NodeManagerHostname': 'jwlogin24.juwels', 'NodeManagerPort': 34927, 'ObjectManagerPort': 41163, 'ObjectStoreSocketName': '/tmp/UID/ray/session_2023-04-05_09-43-27_453188_1518/sockets/plasma_store', 'RayletSocketName': '/tmp/UID/ray/session_2023-04-05_09-43-27_453188_1518/sockets/raylet', 'MetricsExportPort': 60382, 'NodeName': '10.13.22.34', 'alive': True, 'Resources': {'memory': 216577999872.0, 'CPU': 2.0, 'node:10.13.22.34': 1.0, 'object_store_memory': 97104857088.0}}, {'NodeID': 'cf76300b5a3d269db3f859ab1b3f87ce03d8ebd344a37bd7421c2dd0', 'Alive': True, 'NodeManagerAddress': '10.13.31.137', 'NodeManagerHostname': 'jwb0807.juwels', 'NodeManagerPort': 34425, 'ObjectManagerPort': 34533, 'ObjectStoreSocketName': '/tmp/ray/session_2023-04-05_09-43-27_453188_1518/sockets/plasma_store', 'RayletSocketName': '/tmp/ray/session_2023-04-05_09-43-27_453188_1518/sockets/raylet', 'MetricsExportPort': 62643, 'NodeName': '10.13.31.137', 'alive': True, 'Resources': {'accelerator_type:A100': 1.0, 'object_store_memory': 154205725900.0, 'CPU': 96.0, 'GPU': 1.0, 'memory': 359813360436.0, 'node:10.13.31.137': 1.0}}]
{'CPU': 194.0, 'GPU': 5.0, 'accelerator_type:A100': 2.0, 'memory': 935603230926.0, 'object_store_memory': 405258527538.0, 'node:10.13.31.131': 1.0, 'node:10.13.22.34': 1.0, 'node:10.13.31.137': 1.0}

运行的Tuner代码

res = cluster_resources()
num_gpu=int(res["GPU"])
tuner = Tuner(tune.with_resources(
                trainable=fun_to_tune,
                resources = tune.PlacementGroupFactory([{"CPU": 0, "GPU": 1}])),
                param_space = parameters,
                run_config = RunConfig(name, verbose=3,
                                        local_dir= os.path.join(os.environ["SCR"],"ray_results"),
                                        progress_reporter = rep,
                                        failure_config=FailureConfig(fail_fast=True)),
                tune_config = TuneConfig(metric = "mean_JSD_impr",mode = "max",
                                            num_samples = -1, time_budget_s = 3600*4,
                                            trial_dirname_creator = mk_tname,
                                            max_concurrent_trials=num_gpu))
tuner.fit()

当前Ray仅识别到5块GPU,只能启动5个并发trial,需要利用每个节点的全部4块GPU。


解决方案

1. 显式指定工作节点GPU数量

Ray自动检测GPU时可能受Slurm环境限制,启动工作节点时手动指定--num-gpus参数,强制Ray识别全部4块GPU:
修改工作节点启动命令:

srun --gres=gpu:4 --nodes=1 --ntasks=1 -w "\${node_i}" ray start --address 10.13.22.34:6374 --num-gpus 4 --block &

2. 确保CUDA环境变量正确传递

部分场景下Slurm会限制CUDA_VISIBLE_DEVICES的范围,启动Ray前显式设置该变量包含全部4块GPU:

srun --gres=gpu:4 --nodes=1 --ntasks=1 -w "\${node_i}" bash -c 'export CUDA_VISIBLE_DEVICES=0,1,2,3; ray start --address 10.13.22.34:6374 --num-gpus 4 --block &'

3. 验证资源检测结果

启动所有节点后,重新运行ray.nodes()和ray.cluster_resources(),确认每个工作节点的GPU资源值为4.0,集群总GPU数为4*工作节点数。

4. 调整Tuner并发设置

无需手动指定max_concurrent_trials,让Ray自动根据集群资源调度并发任务:

tune_config = TuneConfig(metric = "mean_JSD_impr",mode = "max",
                        num_samples = -1, time_budget_s = 3600*4,
                        trial_dirname_creator = mk_tname)

5. 升级Ray版本

旧版本Ray存在多GPU节点检测bug,建议升级到最新稳定版,避免环境兼容问题。


内容的提问来源于stack exchange,提问作者Wacken0013

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 04:03:09