Slurm集群中Ray分布式HPO的GPU检测异常问题求助
问题
在Slurm集群上进行分布式HPO(超参数优化)时,Ray无法正确检测GPU资源。集群包含仅运行调度器的CPU头节点,以及多台配置相同、各带4块GPU的工作节点,但Ray仅能识别其中一个节点的全部4块GPU,其余节点仅识别1块GPU。
集群启动配置
- 头节点启动命令:
ray start --head --temp-dir /tmp/UID/ray/ --node-ip-address=10.13.22.34 --port=6374 --block --dashboard-port 3232 --dashboard-host 0.0.0.0 --num-cpus 2 &
- 工作节点启动脚本(sbatch):
srun --gres=gpu:4 --nodes=1 --ntasks=1 -w "\${node_i}" ray start --address 10.13.22.34:6374 --block &
Ray资源检测输出
ray.nodes()和ray.cluster_resources()的输出:
[{'NodeID': '3b793f590f23b9a3f8eb4da86f870a2c6a20967ea1ea1093945fb2d9', 'Alive': True, 'NodeManagerAddress': '10.13.31.131', 'NodeManagerHostname': 'jwb0801.juwels', 'NodeManagerPort': 44655, 'ObjectManagerPort': 40891, 'ObjectStoreSocketName': '/tmp/ray/session_2023-04-05_09-43-27_453188_1518/sockets/plasma_store', 'RayletSocketName': '/tmp/ray/session_2023-04-05_09-43-27_453188_1518/sockets/raylet', 'MetricsExportPort': 43462, 'NodeName': '10.13.31.131', 'alive': True, 'Resources': {'CPU': 96.0, 'memory': 359211870618.0, 'GPU': 4.0, 'node:10.13.31.131': 1.0, 'object_store_memory': 153947944550.0, 'accelerator_type:A100': 1.0}}, {'NodeID': '5f803a7dd410074ea9d3fb3058b3ab3044011c54536c7808bd9a947c', 'Alive': True, 'NodeManagerAddress': '10.13.22.34', 'NodeManagerHostname': 'jwlogin24.juwels', 'NodeManagerPort': 34927, 'ObjectManagerPort': 41163, 'ObjectStoreSocketName': '/tmp/UID/ray/session_2023-04-05_09-43-27_453188_1518/sockets/plasma_store', 'RayletSocketName': '/tmp/UID/ray/session_2023-04-05_09-43-27_453188_1518/sockets/raylet', 'MetricsExportPort': 60382, 'NodeName': '10.13.22.34', 'alive': True, 'Resources': {'memory': 216577999872.0, 'CPU': 2.0, 'node:10.13.22.34': 1.0, 'object_store_memory': 97104857088.0}}, {'NodeID': 'cf76300b5a3d269db3f859ab1b3f87ce03d8ebd344a37bd7421c2dd0', 'Alive': True, 'NodeManagerAddress': '10.13.31.137', 'NodeManagerHostname': 'jwb0807.juwels', 'NodeManagerPort': 34425, 'ObjectManagerPort': 34533, 'ObjectStoreSocketName': '/tmp/ray/session_2023-04-05_09-43-27_453188_1518/sockets/plasma_store', 'RayletSocketName': '/tmp/ray/session_2023-04-05_09-43-27_453188_1518/sockets/raylet', 'MetricsExportPort': 62643, 'NodeName': '10.13.31.137', 'alive': True, 'Resources': {'accelerator_type:A100': 1.0, 'object_store_memory': 154205725900.0, 'CPU': 96.0, 'GPU': 1.0, 'memory': 359813360436.0, 'node:10.13.31.137': 1.0}}] {'CPU': 194.0, 'GPU': 5.0, 'accelerator_type:A100': 2.0, 'memory': 935603230926.0, 'object_store_memory': 405258527538.0, 'node:10.13.31.131': 1.0, 'node:10.13.22.34': 1.0, 'node:10.13.31.137': 1.0}
运行的Tuner代码
res = cluster_resources() num_gpu=int(res["GPU"]) tuner = Tuner(tune.with_resources( trainable=fun_to_tune, resources = tune.PlacementGroupFactory([{"CPU": 0, "GPU": 1}])), param_space = parameters, run_config = RunConfig(name, verbose=3, local_dir= os.path.join(os.environ["SCR"],"ray_results"), progress_reporter = rep, failure_config=FailureConfig(fail_fast=True)), tune_config = TuneConfig(metric = "mean_JSD_impr",mode = "max", num_samples = -1, time_budget_s = 3600*4, trial_dirname_creator = mk_tname, max_concurrent_trials=num_gpu)) tuner.fit()
当前Ray仅识别到5块GPU,只能启动5个并发trial,需要利用每个节点的全部4块GPU。
解决方案
1. 显式指定工作节点GPU数量
Ray自动检测GPU时可能受Slurm环境限制,启动工作节点时手动指定--num-gpus参数,强制Ray识别全部4块GPU:
修改工作节点启动命令:
srun --gres=gpu:4 --nodes=1 --ntasks=1 -w "\${node_i}" ray start --address 10.13.22.34:6374 --num-gpus 4 --block &
2. 确保CUDA环境变量正确传递
部分场景下Slurm会限制CUDA_VISIBLE_DEVICES的范围,启动Ray前显式设置该变量包含全部4块GPU:
srun --gres=gpu:4 --nodes=1 --ntasks=1 -w "\${node_i}" bash -c 'export CUDA_VISIBLE_DEVICES=0,1,2,3; ray start --address 10.13.22.34:6374 --num-gpus 4 --block &'
3. 验证资源检测结果
启动所有节点后,重新运行ray.nodes()和ray.cluster_resources(),确认每个工作节点的GPU资源值为4.0,集群总GPU数为4*工作节点数。
4. 调整Tuner并发设置
无需手动指定max_concurrent_trials,让Ray自动根据集群资源调度并发任务:
tune_config = TuneConfig(metric = "mean_JSD_impr",mode = "max", num_samples = -1, time_budget_s = 3600*4, trial_dirname_creator = mk_tname)
5. 升级Ray版本
旧版本Ray存在多GPU节点检测bug,建议升级到最新稳定版,避免环境兼容问题。
内容的提问来源于stack exchange,提问作者Wacken0013
相关产品推荐
相关产品推荐

