构建Vertex AI自定义容器时如何缓存Python包以缩短部署时间?
解决Vertex AI自定义预测容器部署耗时过长的问题
方案1:在基础镜像中预安装大型Python依赖
核心思路是把体积大、更新频率低的机器学习包(如TensorFlow、PyTorch等)直接预装在基础镜像中,避免每次部署重复下载安装。
- 新建
heavy_requirements.txt文件,列出所有不常变更的大型依赖:
tensorflow==2.15.0 torch==2.1.0 scikit-learn==1.3.2
- 修改基础镜像的Dockerfile,加入预安装步骤:
FROM python:3.10 RUN apt update && apt install -y my-depencency # 预装大型依赖,--no-cache-dir可减少镜像体积 COPY heavy_requirements.txt . RUN pip install --no-cache-dir -r heavy_requirements.txt
- 构建并推送该基础镜像到Google Container Registry(GCR):
docker build -t gcr.io/[你的项目ID]/base-with-heavy-deps:v1 . docker push gcr.io/[你的项目ID]/base-with-heavy-deps:v1
- 在
build_cpr_model中使用这个基础镜像,同时将requirements.txt精简为仅包含频繁变更的轻量依赖(如项目自定义库、小型工具包):
from google.cloud.aiplatform.prediction import LocalModel local_model = LocalModel.build_cpr_model( USER_SRC_DIR, IMAGE_NAME, predictor=MyPredictor, requirements_path=os.path.join(USER_SRC_DIR, "light_requirements.txt"), base_image="gcr.io/[你的项目ID]/base-with-heavy-deps:v1", ) local_model.push_image()
这样每次部署仅需安装少量新增依赖,大幅缩短耗时。
方案2:自定义Dockerfile替代build_cpr_model
如果build_cpr_model灵活性不足,可完全手动编写最终镜像的Dockerfile,直接复用基础镜像的所有依赖:
- 编写最终镜像的Dockerfile:
# 基于预装好依赖的基础镜像 FROM gcr.io/[你的项目ID]/base-with-heavy-deps:v1 # 复制自定义预测代码到容器 COPY ./[你的源码目录] /app WORKDIR /app # 仅安装新增依赖(若有) RUN if [ -f requirements.txt ]; then pip install --no-cache-dir -r requirements.txt; fi # 设置Vertex AI预测服务所需的环境变量和启动命令 ENV AIP_HTTP_PORT=8080 CMD ["python", "-m", "predictor_server"]
- 手动构建并推送镜像:
docker build -t gcr.io/[你的项目ID]/final-prediction-image:v1 . docker push gcr.io/[你的项目ID]/final-prediction-image:v1
- 通过Vertex AI API直接导入镜像创建模型:
from google.cloud import aiplatform aiplatform.init(project="[你的项目ID]", region="[你的区域]") model = aiplatform.Model.upload( display_name="custom-prediction-model", serving_container_image_uri="gcr.io/[你的项目ID]/final-prediction-image:v1", serving_container_predict_route="/predict", serving_container_health_route="/health" )
这种方式完全掌控镜像构建流程,确保基础镜像的依赖100%被复用。
方案3:优化build_cpr_model的依赖安装逻辑
pip默认会跳过已安装且版本匹配的包,只需确保requirements.txt中依赖的版本与基础镜像预安装的版本完全一致,build_cpr_model构建时就不会重复下载安装这些包:
- 例如基础镜像安装了
tensorflow==2.15.0,则requirements.txt中同样写tensorflow==2.15.0,pip会直接跳过该包; - 避免使用无版本限制的依赖(如
tensorflow),否则pip可能尝试更新包,导致重复下载。
内容的提问来源于stack exchange,提问作者Radu Dilirici
相关产品推荐
相关产品推荐

