You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在GCP Vertex AI中部署自定义模型至私有服务端点

自定义TensorFlow模型部署到GCP Vertex AI私有端点分步指南

前置准备

  • 已安装并初始化 gcloud CLI,且切换到目标项目:gcloud config set project YOUR_PROJECT_ID
  • TensorFlow模型已导出为SavedModel格式(目录需包含saved_model.pb和variables子目录)
  • 拥有GCP项目的Editor或Vertex AI Admin权限
  • 已启用Vertex AI、Compute Engine、VPC Network API

方案一:基于Private Services Access部署

步骤1:配置VPC与专用连接

  • 启用目标VPC的私有Google访问:
    gcloud compute networks update YOUR_VPC_NAME --enable-private-google-access
    
  • 创建专用访问IP范围并建立对等连接:
    1. 预留一个不冲突的/24私有IP段(如10.0.0.0/24):
      gcloud compute addresses create vertex-ai-range --global --prefix-length=24 --network=YOUR_VPC_NAME --purpose=VPC_PEERING
      
    2. 建立与Vertex AI服务的VPC对等连接:
      gcloud services vpc-peerings connect --service=servicenetworking.googleapis.com --ranges=vertex-ai-range --network=YOUR_VPC_NAME --project=YOUR_PROJECT_ID
      

步骤2:上传模型到GCS

将本地SavedModel目录上传至Google Cloud Storage:

gsutil cp -r YOUR_LOCAL_MODEL_DIR gs://YOUR_BUCKET_NAME/models/tf-model/

步骤3:创建Vertex AI模型资源

gcloud ai models upload --region=YOUR_REGION --display-name=tf-private-model --artifact-uri=gs://YOUR_BUCKET_NAME/models/tf-model/ --framework=TENSORFLOW

步骤4:创建私有端点

gcloud ai endpoints create --region=YOUR_REGION --display-name=tf-private-endpoint --network=YOUR_VPC_NAME

步骤5:部署模型到端点

指定机器类型与副本数完成部署:

gcloud ai endpoints deploy-model YOUR_ENDPOINT_ID --region=YOUR_REGION --model=YOUR_MODEL_ID --machine-type=n1-standard-4 --min-replica-count=1 --max-replica-count=3

方案二:基于Private Service Connect部署

步骤1:配置VPC与PSC端点

  • 创建或使用现有VPC子网(如10.1.0.0/24),创建PSC转发规则:
    gcloud compute forwarding-rules create psc-vertex-endpoint --region=YOUR_REGION --network=YOUR_VPC_NAME --subnet=YOUR_SUBNET_NAME --address=YOUR_INTERNAL_IP --target-http-proxy=vertex-proxy --psc-target-service=servicenetworking.googleapis.com/projects/YOUR_PROJECT_ID/regions/YOUR_REGION/services/vertex-ai
    

    注:YOUR_INTERNAL_IP需为子网内未占用的私有IP

步骤2:上传模型与创建模型资源

同方案一的步骤2、3

步骤3:创建PSC关联的私有端点

gcloud ai endpoints create --region=YOUR_REGION --display-name=tf-psc-endpoint --enable-private-service-connect --psc-allowed-projects=YOUR_PROJECT_ID

步骤4:部署模型到PSC端点

同方案一的步骤5


验证端点可用性

在VPC内部的虚拟机上执行预测测试:

gcloud ai endpoints predict YOUR_ENDPOINT_ID --region=YOUR_REGION --json-request=test-input.json

test-input.json需匹配模型输入格式,示例:{"instances": [[1.0, 2.0, 3.0]]}

注意事项

  • 确保VPC防火墙规则允许内部流量访问端点默认端口8080
  • 自定义服务账号需具备aiplatform.serviceAgent角色或对应权限
  • 私有端点仅接受VPC内部或对等VPC的流量,公网无法直接访问

内容的提问来源于stack exchange,提问作者A.R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 18:42:53