如何在GCP Vertex AI中部署自定义模型至私有服务端点
自定义TensorFlow模型部署到GCP Vertex AI私有端点分步指南
前置准备
- 已安装并初始化
gcloudCLI,且切换到目标项目:gcloud config set project YOUR_PROJECT_ID - TensorFlow模型已导出为SavedModel格式(目录需包含
saved_model.pb和variables子目录) - 拥有GCP项目的Editor或Vertex AI Admin权限
- 已启用Vertex AI、Compute Engine、VPC Network API
方案一:基于Private Services Access部署
步骤1:配置VPC与专用连接
- 启用目标VPC的私有Google访问:
gcloud compute networks update YOUR_VPC_NAME --enable-private-google-access - 创建专用访问IP范围并建立对等连接:
- 预留一个不冲突的/24私有IP段(如
10.0.0.0/24):gcloud compute addresses create vertex-ai-range --global --prefix-length=24 --network=YOUR_VPC_NAME --purpose=VPC_PEERING - 建立与Vertex AI服务的VPC对等连接:
gcloud services vpc-peerings connect --service=servicenetworking.googleapis.com --ranges=vertex-ai-range --network=YOUR_VPC_NAME --project=YOUR_PROJECT_ID
- 预留一个不冲突的/24私有IP段(如
步骤2:上传模型到GCS
将本地SavedModel目录上传至Google Cloud Storage:
gsutil cp -r YOUR_LOCAL_MODEL_DIR gs://YOUR_BUCKET_NAME/models/tf-model/
步骤3:创建Vertex AI模型资源
gcloud ai models upload --region=YOUR_REGION --display-name=tf-private-model --artifact-uri=gs://YOUR_BUCKET_NAME/models/tf-model/ --framework=TENSORFLOW
步骤4:创建私有端点
gcloud ai endpoints create --region=YOUR_REGION --display-name=tf-private-endpoint --network=YOUR_VPC_NAME
步骤5:部署模型到端点
指定机器类型与副本数完成部署:
gcloud ai endpoints deploy-model YOUR_ENDPOINT_ID --region=YOUR_REGION --model=YOUR_MODEL_ID --machine-type=n1-standard-4 --min-replica-count=1 --max-replica-count=3
方案二:基于Private Service Connect部署
步骤1:配置VPC与PSC端点
- 创建或使用现有VPC子网(如
10.1.0.0/24),创建PSC转发规则:gcloud compute forwarding-rules create psc-vertex-endpoint --region=YOUR_REGION --network=YOUR_VPC_NAME --subnet=YOUR_SUBNET_NAME --address=YOUR_INTERNAL_IP --target-http-proxy=vertex-proxy --psc-target-service=servicenetworking.googleapis.com/projects/YOUR_PROJECT_ID/regions/YOUR_REGION/services/vertex-ai注:YOUR_INTERNAL_IP需为子网内未占用的私有IP
步骤2:上传模型与创建模型资源
同方案一的步骤2、3
步骤3:创建PSC关联的私有端点
gcloud ai endpoints create --region=YOUR_REGION --display-name=tf-psc-endpoint --enable-private-service-connect --psc-allowed-projects=YOUR_PROJECT_ID
步骤4:部署模型到PSC端点
同方案一的步骤5
验证端点可用性
在VPC内部的虚拟机上执行预测测试:
gcloud ai endpoints predict YOUR_ENDPOINT_ID --region=YOUR_REGION --json-request=test-input.json
test-input.json需匹配模型输入格式,示例:
{"instances": [[1.0, 2.0, 3.0]]}
注意事项
- 确保VPC防火墙规则允许内部流量访问端点默认端口8080
- 自定义服务账号需具备
aiplatform.serviceAgent角色或对应权限 - 私有端点仅接受VPC内部或对等VPC的流量,公网无法直接访问
内容的提问来源于stack exchange,提问作者A.R
相关产品推荐
相关产品推荐

