添加torch依赖后无法在GCP App Engine Flex部署求助
GCP App Engine Flexible部署Torch时的构建超时与磁盘空间不足问题
问题背景
- 执行
poetry add torch添加依赖后,pyproject.toml生成配置:torch = "^2.0.0" - 部署到GCP App Engine Flexible时出现两个核心问题:
- 构建时长超40分钟,需临时延长超时时间
- 抛出磁盘空间不足错误:
ERROR: (gcloud.app.deploy) Error Response: [9] An internal error occurred while processing task /app-engine-flex/flex_await_healthy/flex_await_healthy>2023-04-06T18:18:10.984Z2529.ue.0: No sufficient free disk space left for your App Engine Flexible application. Please increase your VM instances disk size in the resource settings in the app.yaml file for your deployment and retry. See https://cloud.google.com/appengine/docs/flexible/python/reference/app-yaml#resource-settings for how to set disk resource.
- 已调整
app.yaml资源配置:resources: cpu: 3 memory_gb: 15 #cpu * [1.0 - 6.5] - 0.4 , we are using 3 for the range value disk_size_gb: 40 - 当前临时用numpy替代Torch完成计算,需解决Torch的部署问题
解决方案
1. 优化Torch依赖安装,避免源码编译
Torch默认安装可能触发源码编译,既耗时又占用大量磁盘,建议直接安装预编译wheel包:
- 安装CPU-only版本(无需CUDA支持时优先选择):
执行命令:
对应的poetry add torch==2.0.0+cpu --extra-index-url https://download.pytorch.org/whl/cpupyproject.toml会自动更新为:torch = { version = "2.0.0+cpu", url = "https://download.pytorch.org/whl/cpu/torch-2.0.0%2Bcpu-cp310-cp310-linux_x86_64.whl" } - 指定CUDA版本预编译包(需要GPU加速时):
以CUDA 11.7为例,执行:poetry add torch==2.0.0+cu117 --extra-index-url https://download.pytorch.org/whl/cu117
2. 优化GCP构建流程
- 启用构建缓存:
创建或修改cloudbuild.yaml,配置缓存复用依赖安装层,减少重复安装耗时:steps: - name: 'python:3.10' args: ['pip', 'install', '--upgrade', 'pip', 'poetry'] cache: paths: - /root/.cache/pip/** - name: 'python:3.10' args: ['poetry', 'install', '--no-root'] cache: paths: - .venv/** - 使用轻量级基础镜像:
在app.yaml中指定精简Python镜像,降低磁盘占用:runtime: python env: flex runtime_config: python_version: 3.10 image: python:3.10-slim
3. 排查磁盘占用细节
- 在构建步骤中添加磁盘检查命令,定位占用大户:
在cloudbuild.yaml中增加步骤:
查看输出确认哪个目录占用过多磁盘,针对性清理(比如临时编译文件、冗余依赖)- name: 'python:3.10' args: ['df', '-h']
内容的提问来源于stack exchange,提问作者Sandeep Kumar Pani
相关产品推荐
相关产品推荐

