Vertex AI自定义容器/predict端点返回405 Method Not Allowed错误求助
Vertex AI自定义容器部署405 Method Not Allowed问题排查与解决
问题描述
在Vertex AI部署自定义容器时,本地基于Gunicorn运行的Flask服务器中,/predict和/health端点均正常响应,但Vertex AI调用预测API时始终返回405 Method Not Allowed错误。
配置信息
- 容器:自定义Docker容器,暴露8080端口
- 模型上传参数:使用以下标记上传模型到Vertex AI:
--container-predict-route=/predict--container-health-route=/health
- 预测调用:通过Google Cloud AI Platform客户端库调用预测API
观察结果
- Vertex AI PredictionService发送请求至类似
/v1/endpoints/<ENDPOINT_ID>/deployedModels/<DEPLOYED_MODEL_ID>:predict的URL,但服务器返回405 - 终端发起GET请求到端点可获有效响应,但调用
/predict或/rawPredict仍返回405 - 服务器持续运行,每10秒有日志:
<GET /v1/endpoints/<ENDPOINT>/deployedModels/<MODEL> HTTP/1.1" 200 OK> - 已添加多条路由(包括通配路由)处理上述URL,但错误仍存在
相关代码
Dockerfile
FROM nvidia/cuda:12.2.0-runtime-ubuntu20.04 RUN apt-get update && apt-get install -y --no-install-recommends \ wget \ curl \ python3-dev \ python3-pip \ python3-setuptools && \ rm -rf /var/lib/apt/lists/* RUN ln -sf /usr/bin/python3 /usr/bin/python WORKDIR /app COPY requirements.txt . RUN pip install --no-cache-dir --upgrade pip && \ pip install --no-cache-dir torch>=1.12.0 torchvision>=0.13.0 && \ if [ -f requirements.txt ]; then pip install --no-cache-dir -r requirements.txt; fi COPY . . EXPOSE 8080 CMD ["gunicorn", "-w", "1", "-b", "0.0.0.0:8080", "main:app"]
Vertex AI API调用函数(call_vertex_ai)
def call_vertex_ai(gcs_uri: str, additional_args: dict): client_options = {"api_endpoint": f"{REGION}-aiplatform.googleapis.com"} client = aiplatform.gapic.PredictionServiceClient( client_options=client_options) instance = predict.instance.ImageClassificationPredictionInstance( content=gcs_uri # GCS path for image file ).to_value() instances = [instance] parameters = predict.params.ImageClassificationPredictionParams( confidence_threshold=additional_args.get("threshold", 0.5), ).to_value() endpoint = client.endpoint_path( project=PROJECT_ID, location=REGION, endpoint=ENDPOINT_ID ) response = client.predict( endpoint=endpoint, instances=instances, parameters=parameters) return response.predictions
Flask服务代码(main.py)
... some imports ... app = Flask(__name__) def load_model(): ... load_model() def handle_predict(): ... code ... detections = [{ "bbox": bbox.tolist() if isinstance(bbox, np.ndarray) else bbox, "class": class_name, "score": float(score), } for bbox, class_name, score in zip(draw_boxes, pred_classes, scores)] return jsonify({"predictions": detections}) @app.post("/predict") def predict(): return handle_predict() @app.route("/health", methods=["GET"]) def health(): return jsonify({"status": "healthy"}) @app.route("/v1/endpoints/<endpoint_id>/deployedModels/<path:deployed_model_path>", methods=["POST"]) def predict_deployed_model(endpoint_id, deployed_model_path): if not deployed_model_path.endswith(":predict"): return "Not Found", 404 return handle_predict() @app.route("/v1/endpoints/<endpoint_id>/deployedModels/<deployed_model_id>:predict", methods=["POST"]) def predict_deployed_model_direct(endpoint_id, deployed_model_id): return handle_predict() @app.route("/v1/endpoints/<endpoint_id>/deployedModels/<deployed_model_id>:rawPredict", methods=["POST"]) def raw_predict_deployed_model(endpoint_id, deployed_model_id): return handle_predict() @app.before_request def log_request_info(): logger.info(f"Received request: {request.method} {request.url}") logger.info(f"Headers: {dict(request.headers)}") logger.info(f"Body: {request.get_data().decode('utf-8')}")
部署脚本(deploy.sh)
gcloud builds submit \ --tag "${IMAGE_NAME}:latest" \ --gcs-source-staging-dir="gs://$BUCKET_NAME/source" \ --gcs-log-dir="gs://$BUCKET_NAME/logs" LATEST_IMAGE="${IMAGE_NAME}:latest" gcloud ai models upload \ --region="${REGION}" \ --display-name="weldpredict-model" \ --container-image-uri="${LATEST_IMAGE}" \ --container-ports=8080 \ --container-predict-route=/predict \ --container-health-route=/health ENDPOINT_ID=$(gcloud ai endpoints list --region="${REGION}" --format="value(ENDPOINT_ID)") DEPLOYED_MODEL_ID=$(gcloud ai endpoints describe "${ENDPOINT_ID}" --region="${REGION}" --format="value(deployedModels.id)") gcloud ai endpoints undeploy-model "${ENDPOINT_ID}" --deployed-model-id="${DEPLOYED_MODEL_ID}" --region="${REGION}" --quiet gcloud ai endpoints deploy-model "${ENDPOINT_ID}" \ --model="${MODEL_ID}" \ --region="${REGION}" \ --display-name="weldpredict-deployment" \ --machine-type=n1-standard-4 \ --accelerator=type=nvidia-tesla-t4,count=1 \ --min-replica-count=1 \ --max-replica-count=1 \ --traffic-split=0=100
问题总结
- 核心问题:调用Vertex AI预测时返回
405 Method Not Allowed错误 - 矛盾点:本地服务正常,平台调用
/predict或/rawPredict返回405 - 配置情况:已指定
--container-predict-route=/predict,但Vertex AI发送的请求未匹配路由 - 尝试方案:添加多条路由仍无法解决
原因分析与解决方案
1. 路由方法与版本兼容问题
- 排查点:检查容器内Flask版本,
@app.post是Flask 2.0+的语法,若版本过低会导致路由无法识别POST请求。 - 修复方案:将
@app.post("/predict")改为兼容写法:
增加OPTIONS方法支持,避免预检请求触发405。@app.route("/predict", methods=["POST", "OPTIONS"]) def predict(): if request.method == "OPTIONS": return "", 200 return handle_predict()
2. Vertex AI请求转发参数缺失
- 排查点:模型上传时指定的
--container-predict-route参数,在部署到端点时可能未被继承,导致平台未将请求转发到/predict。 - 修复方案:在部署端点时重新指定容器路由参数:
gcloud ai endpoints deploy-model "${ENDPOINT_ID}" \ --model="${MODEL_ID}" \ --region="${REGION}" \ --display-name="weldpredict-deployment" \ --machine-type=n1-standard-4 \ --accelerator=type=nvidia-tesla-t4,count=1 \ --min-replica-count=1 \ --max-replica-count=1 \ --traffic-split=0=100 \ --container-predict-route=/predict \ --container-health-route=/health
3. 路由匹配优先级与规则问题
- 排查点:Flask路由按定义顺序匹配,通配路由可能覆盖了具体路由;或路径中的冒号导致路由变量匹配异常。
- 修复方案:
- 调整路由顺序,将具体路由(如
:predict、:rawPredict)放在通配路由之前; - 临时添加全局通配路由,验证请求是否能被捕获:
若该路由能接收请求,说明原路由匹配规则存在问题,需优化路径匹配逻辑。@app.route("/", defaults={"path": ""}, methods=["POST"]) @app.route("/<path:path>", methods=["POST"]) def catch_all(path): logger.info(f"Catch all request path: {path}") return handle_predict()
- 调整路由顺序,将具体路由(如
4. 容器网络与Gunicorn配置验证
- 排查点:确认Gunicorn绑定地址为
0.0.0.0:8080,容器端口暴露正常,无网络策略阻止请求;检查容器内是否有其他进程占用8080端口。 - 修复方案:在Dockerfile中增加端口检查命令,或在启动Gunicorn时添加日志参数(如
--access-logfile -),验证请求是否到达容器。
内容的提问来源于stack exchange,提问作者Giulio Manuzzi
相关产品推荐
相关产品推荐

