You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Vertex AI用VLLM部署Llama3 70B时遭遇Ray磁盘空间错误

Vertex AI部署vLLM运行Llama-3-70B-Instruct时/tmp/ray空间不足的解决方法
  • 调整Ray临时目录到大容量磁盘
    Vertex AI预测容器默认的/tmp磁盘空间有限,Ray默认用这个目录存临时数据。可以把Ray临时目录改到Vertex挂载的SSD路径/mnt/disks/ssd0:

    • 容器构建时提前创建目录并开放权限:
      RUN mkdir -p /mnt/disks/ssd0/ray_tmp && chmod 777 /mnt/disks/ssd0/ray_tmp
      
    • 在vLLM初始化前设置环境变量:
      import os
      os.environ["RAY_TMPDIR"] = "/mnt/disks/ssd0/ray_tmp"
      
  • 增大共享内存配置
    当前设置的16GB共享内存对于8卡L4运行70B模型不够,建议上调到64GB:

    model_resource = aiplatform.Model.upload(
        serving_container_image_uri=serving_container_image_uri,
        serving_container_shared_memory_size_mb=65536,
        # 其他参数保留
    )
    
  • 优化vLLM的Ray内存配置
    限制Ray对象存储的内存使用,避免触发磁盘溢出检查:

    self.model = LLM(
        model=model_config.model_hf_name,
        dtype="auto",
        tensor_parallel_size=model_config.tensor_parallel_size,
        enforce_eager=model_config.enforce_eager,
        disable_custom_all_reduce=model_config.disable_custom_all_reduce,
        worker_use_ray=bool(model_config.tensor_parallel_size > 1),
        enable_prefix_caching=False,
        max_model_len=model_config.max_seq_len,
        ray_args={
            "object_store_memory": 20 * 1024**3,  # 20GB,可根据实际调整
            "memory": 30 * 1024**3,
        }
    )
    

    也可以直接禁用Ray的内存监控,添加环境变量RAY_DISABLE_MEMORY_MONITOR=1。

  • 降低vLLM内存占用

    • 改用bfloat16 dtype(L4原生支持,比auto更省内存):
      self.model = LLM(
          model=model_config.model_hf_name,
          dtype="bfloat16",
          # 其他参数保留
      )
      
    • 开启enable_chunked_prefill=True,削减预填充阶段的内存峰值。
  • 验证磁盘挂载情况
    在容器启动脚本里加df -h命令,确认/mnt/disks/ssd0的可用空间足够(L4实例SSD至少有300GB)。

内容的提问来源于stack exchange,提问作者Tsvi Sabo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 07:43:31