You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ray 2.5遇OOM异常终止整个实验的相关技术咨询

Ray 2.5 单节点调优实验内存问题及疑问

问题背景

本地单节点运行Ray调优实验,共N个试验,同一时间仅运行一个,试验内存需求随超参数变化。Ray 2.5版本中,实验偶尔因ray.exceptions.OutOfMemoryError崩溃,且整个实验直接终止;但Ray 1.13版本仅标记对应试验错误,实验会继续运行。

错误信息

主进程崩溃日志

Traceback (most recent call last):
  File "/home/.../lib/python3.8/site-packages/ray/air/execution/_internal/event_manager.py", line 110, in resolve_future
    result = ray.get(future)
  File "/home/.../lib/python3.8/site-packages/ray/_private/auto_init_hook.py", line 18, in auto_init_wrapper
    return fn(*args, **kwargs)
  File "/home/.../lib/python3.8/site-packages/ray/_private/client_mode_hook.py", line 103, in wrapper
    return func(*args, **kwargs)
  File "/home/.../lib/python3.8/site-packages/ray/_private/worker.py", line 2542, in get
    raise value
ray.exceptions.OutOfMemoryError: Task was killed due to the node running low on memory.
Memory on the node (IP: xxx, ID: 3509da226e9a51e1d3be7f9855d12771eb10a1eea05bf5512de91f73) where the task (actor ID: 91b7d76827552756b57fc3a301000000, name=ImplicitFunc.__init__, pid=134132, memory used=57.48GB) was running was 62.04GB / 62.63GB (0.990643), which exceeds the memory usage threshold of 0.99. Ray killed this worker (ID: a15d6fa9dd1fd4ada92d64a0197f5d7c5cf5731eb8eac96c38d37ced) because it was the most recently scheduled task; to see more information about memory usage on this node, use `ray logs raylet.out -ip xxx`. To see the logs of the worker, use `ray logs worker-a15d6fa9dd1fd4ada92d64a0197f5d7c5cf5731eb8eac96c38d37ced*out -ip xxx.

试验目录内错误日志

Failure # 1 (occurred at 2023-06-21_08-16-56)
The actor died unexpectedly before finishing this task.
    class_name: ImplicitFunc
    actor_id: 6b89355032ae9f5fcad2ced601000000
    pid: 991990
    namespace: 12c53ee2-d178-4697-97a2-d87579e2ef12
    ip: xxxxxx
The actor is dead because its worker process has died. Worker exit type: SYSTEM_ERROR Worker exit detail: Worker unexpectedly exits with a connection error code 2. End of file. There are some potential root causes. (1) The process is killed by SIGKILL by OOM killer due to high memory usage. (2) ray stop --force is called. (3) The worker is crashed unexpectedly due to SIGSEGV or other unexpected errors.

调优代码

tune_config = tune.TuneConfig(num_samples=args.num_samples,
                              metric="AUC",
                              mode="max")
run_config = air.RunConfig(name=args.name,
                           storage_path=args.save_dir,
                           verbose=3)
trainable_with_resources = tune.with_resources(train_validate_model, {"cpu": 24, "gpu": 1})
tuner = tune.Tuner(trainable_with_resources, param_space=config, tune_config=tune_config, run_config=run_config)
tuner.fit()

疑问及解答

1. Ray 1.13与2.5版本异常处理差异原因?

Ray 2.x重构了Air/Tune的执行引擎,默认将OOM这类任务失败视为致命错误,会终止整个调优实验;而Ray 1.x的Tune执行逻辑更宽松,仅标记单个试验失败,实验继续运行。

要恢复1.x的行为,可在RunConfig中配置failure_config,设置fail_fast=False,同时调整max_failures参数允许一定数量的试验失败:

run_config = air.RunConfig(
    name=args.name,
    storage_path=args.save_dir,
    verbose=3,
    failure_config=air.FailureConfig(fail_fast=False, max_failures=10)  # 根据需求设置max_failures
)

2. 内存未清理、残留进程问题如何解决?

  • 内存泄漏排查:试验重启后可正常运行,说明当前试验的内存对象未被正确回收。检查train_validate_model函数内是否存在全局变量、未关闭的文件/句柄、大型张量未及时释放(如PyTorch/TensorFlow模型未调用.to('cpu')或del后未触发GC)。可在试验结束时手动调用gc.collect()强制回收内存。
  • 残留进程处理:实验终止后残留ray::ImplicitFunc进程,是因为Ray的actor进程未被正确销毁。可在脚本退出时添加显式清理逻辑:
    import ray
    # 在tuner.fit()之后或异常捕获块中添加
    ray.shutdown()
    
    若仍有残留,可手动执行命令强制终止Ray相关进程:
    ray stop --force
    pkill -f "ray::ImplicitFunc"
    

3. Ray术语解释及max_restarts/max_task_retries的作用?

  • 术语定义:
    • Task:Ray中最小的执行单元,是无状态的函数调用,执行完成后释放资源。你的train_validate_model如果是普通函数,默认会作为Task执行;但如果它内部创建了Actor(如ImplicitFunc),则会生成Actor实例。
    • Actor:有状态的服务/对象,会长期占用资源,直到被显式销毁。ImplicitFunc就是一个Actor类,实例化后会在Worker进程中运行。
    • Worker:Ray的工作进程,负责执行Task或托管Actor。每个Worker对应一个OS进程,内存占用就是该进程的内存使用量。
  • 参数作用:
    • max_restarts:针对Actor的配置,设置Actor意外崩溃后的最大重启次数。如果你的试验依赖ImplicitFunc Actor,设置该参数可让Actor崩溃后自动重启,继续完成试验。
    • max_task_retries:针对Task的配置,设置Task失败后的重试次数。如果你的train_validate_model是Task,设置该参数可让OOM失败的Task自动重试。
      这两个参数需要结合FailureConfig一起使用,才能让试验失败后自动恢复,而不是终止整个实验。

内容的提问来源于stack exchange,提问作者user1809923

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 16:47:01