You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Vertex AI自定义容器/predict端点返回405 Method Not Allowed错误求助

Vertex AI自定义容器部署405 Method Not Allowed问题排查与解决

问题描述

在Vertex AI部署自定义容器时,本地基于Gunicorn运行的Flask服务器中,/predict和/health端点均正常响应,但Vertex AI调用预测API时始终返回405 Method Not Allowed错误。

配置信息

  • 容器:自定义Docker容器,暴露8080端口
  • 模型上传参数:使用以下标记上传模型到Vertex AI:
    • --container-predict-route=/predict
    • --container-health-route=/health
  • 预测调用:通过Google Cloud AI Platform客户端库调用预测API

观察结果

  • Vertex AI PredictionService发送请求至类似/v1/endpoints/<ENDPOINT_ID>/deployedModels/<DEPLOYED_MODEL_ID>:predict的URL,但服务器返回405
  • 终端发起GET请求到端点可获有效响应,但调用/predict或/rawPredict仍返回405
  • 服务器持续运行,每10秒有日志:<GET /v1/endpoints/<ENDPOINT>/deployedModels/<MODEL> HTTP/1.1" 200 OK>
  • 已添加多条路由(包括通配路由)处理上述URL,但错误仍存在

相关代码

Dockerfile

FROM nvidia/cuda:12.2.0-runtime-ubuntu20.04
RUN apt-get update && apt-get install -y --no-install-recommends \
    wget \
    curl \
    python3-dev \
    python3-pip \
    python3-setuptools && \
    rm -rf /var/lib/apt/lists/*
RUN ln -sf /usr/bin/python3 /usr/bin/python
WORKDIR /app

COPY requirements.txt .
RUN pip install --no-cache-dir --upgrade pip && \
    pip install --no-cache-dir torch>=1.12.0 torchvision>=0.13.0 && \
    if [ -f requirements.txt ]; then pip install --no-cache-dir -r requirements.txt; fi
COPY . .
EXPOSE 8080
CMD ["gunicorn", "-w", "1", "-b", "0.0.0.0:8080", "main:app"]

Vertex AI API调用函数(call_vertex_ai)

def call_vertex_ai(gcs_uri: str, additional_args: dict):
    client_options = {"api_endpoint": f"{REGION}-aiplatform.googleapis.com"}
    client = aiplatform.gapic.PredictionServiceClient(
        client_options=client_options)

    instance = predict.instance.ImageClassificationPredictionInstance(
        content=gcs_uri  # GCS path for image file
    ).to_value()
    instances = [instance]

    parameters = predict.params.ImageClassificationPredictionParams(
        confidence_threshold=additional_args.get("threshold", 0.5),
    ).to_value()

    endpoint = client.endpoint_path(
        project=PROJECT_ID, location=REGION, endpoint=ENDPOINT_ID
    )

    response = client.predict(
        endpoint=endpoint, instances=instances, parameters=parameters)

    return response.predictions

Flask服务代码(main.py)

... some imports ... 
app = Flask(__name__)

def load_model():
    ...

load_model()


def handle_predict():    ... code ...
    detections = [{
        "bbox": bbox.tolist() if isinstance(bbox, np.ndarray) else bbox,
        "class": class_name,
        "score": float(score),
    } for bbox, class_name, score in zip(draw_boxes, pred_classes, scores)]

    return jsonify({"predictions": detections})

@app.post("/predict")
def predict():
    return handle_predict()

@app.route("/health", methods=["GET"])
def health():
    return jsonify({"status": "healthy"})


@app.route("/v1/endpoints/<endpoint_id>/deployedModels/<path:deployed_model_path>", methods=["POST"])
def predict_deployed_model(endpoint_id, deployed_model_path):
    if not deployed_model_path.endswith(":predict"):
        return "Not Found", 404
    return handle_predict()

@app.route("/v1/endpoints/<endpoint_id>/deployedModels/<deployed_model_id>:predict", methods=["POST"])
def predict_deployed_model_direct(endpoint_id, deployed_model_id):
    return handle_predict()

@app.route("/v1/endpoints/<endpoint_id>/deployedModels/<deployed_model_id>:rawPredict", methods=["POST"])
def raw_predict_deployed_model(endpoint_id, deployed_model_id):
    return handle_predict()

@app.before_request
def log_request_info():
    logger.info(f"Received request: {request.method} {request.url}")
    logger.info(f"Headers: {dict(request.headers)}")
    logger.info(f"Body: {request.get_data().decode('utf-8')}")

部署脚本(deploy.sh)

gcloud builds submit \
  --tag "${IMAGE_NAME}:latest" \
  --gcs-source-staging-dir="gs://$BUCKET_NAME/source" \
  --gcs-log-dir="gs://$BUCKET_NAME/logs"

LATEST_IMAGE="${IMAGE_NAME}:latest"

gcloud ai models upload \
  --region="${REGION}" \
  --display-name="weldpredict-model" \
  --container-image-uri="${LATEST_IMAGE}" \
  --container-ports=8080 \
  --container-predict-route=/predict \
  --container-health-route=/health

ENDPOINT_ID=$(gcloud ai endpoints list --region="${REGION}" --format="value(ENDPOINT_ID)")

DEPLOYED_MODEL_ID=$(gcloud ai endpoints describe "${ENDPOINT_ID}" --region="${REGION}" --format="value(deployedModels.id)")
gcloud ai endpoints undeploy-model "${ENDPOINT_ID}" --deployed-model-id="${DEPLOYED_MODEL_ID}" --region="${REGION}" --quiet

gcloud ai endpoints deploy-model "${ENDPOINT_ID}" \
    --model="${MODEL_ID}" \
    --region="${REGION}" \
    --display-name="weldpredict-deployment" \
    --machine-type=n1-standard-4 \
    --accelerator=type=nvidia-tesla-t4,count=1 \
    --min-replica-count=1 \
    --max-replica-count=1 \
    --traffic-split=0=100

问题总结

  • 核心问题:调用Vertex AI预测时返回405 Method Not Allowed错误
  • 矛盾点:本地服务正常,平台调用/predict或/rawPredict返回405
  • 配置情况:已指定--container-predict-route=/predict,但Vertex AI发送的请求未匹配路由
  • 尝试方案:添加多条路由仍无法解决

原因分析与解决方案

1. 路由方法与版本兼容问题

  • 排查点:检查容器内Flask版本,@app.post是Flask 2.0+的语法,若版本过低会导致路由无法识别POST请求。
  • 修复方案:将@app.post("/predict")改为兼容写法:
    @app.route("/predict", methods=["POST", "OPTIONS"])
    def predict():
        if request.method == "OPTIONS":
            return "", 200
        return handle_predict()
    
    增加OPTIONS方法支持,避免预检请求触发405。

2. Vertex AI请求转发参数缺失

  • 排查点:模型上传时指定的--container-predict-route参数,在部署到端点时可能未被继承,导致平台未将请求转发到/predict。
  • 修复方案:在部署端点时重新指定容器路由参数:
    gcloud ai endpoints deploy-model "${ENDPOINT_ID}" \
        --model="${MODEL_ID}" \
        --region="${REGION}" \
        --display-name="weldpredict-deployment" \
        --machine-type=n1-standard-4 \
        --accelerator=type=nvidia-tesla-t4,count=1 \
        --min-replica-count=1 \
        --max-replica-count=1 \
        --traffic-split=0=100 \
        --container-predict-route=/predict \
        --container-health-route=/health
    

3. 路由匹配优先级与规则问题

  • 排查点:Flask路由按定义顺序匹配,通配路由可能覆盖了具体路由;或路径中的冒号导致路由变量匹配异常。
  • 修复方案:
    1. 调整路由顺序,将具体路由(如:predict、:rawPredict)放在通配路由之前;
    2. 临时添加全局通配路由,验证请求是否能被捕获:
      @app.route("/", defaults={"path": ""}, methods=["POST"])
      @app.route("/<path:path>", methods=["POST"])
      def catch_all(path):
          logger.info(f"Catch all request path: {path}")
          return handle_predict()
      
      若该路由能接收请求,说明原路由匹配规则存在问题,需优化路径匹配逻辑。

4. 容器网络与Gunicorn配置验证

  • 排查点:确认Gunicorn绑定地址为0.0.0.0:8080,容器端口暴露正常,无网络策略阻止请求;检查容器内是否有其他进程占用8080端口。
  • 修复方案:在Dockerfile中增加端口检查命令,或在启动Gunicorn时添加日志参数(如--access-logfile -),验证请求是否到达容器。

内容的提问来源于stack exchange,提问作者Giulio Manuzzi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 02:39:54