You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Vertex AI端点部署自定义容器后查询报500内部服务器错误排查

Vertex AI自定义容器端点500内部服务器错误排查

问题描述

已成功将搭载PaLM大语言模型的自定义容器部署到Vertex AI端点,但通过Vertex AI API或gcloud cli查询端点时,收到500 Internal Server Error回复。

错误信息(翻译后)

500 Internal Server Error
无法获取预测结果。请检查端点日志以获取详细信息。

可能的原因分析

  • 请求格式解析错误:Vertex AI端点发送的预测请求结构为{"instances": [{"question": "..."}]},但当前代码直接读取body["question"],会因找不到对应字段引发KeyError,导致500错误。
  • 资源重复初始化:每次预测请求都重新创建MatchingEngine和TextGenerationModel实例,会增加资源开销,甚至引发初始化超时。
  • 服务账号权限不足:容器内代码调用Vertex AI的Matching Engine、文本生成模型及GCS桶时,部署端点使用的服务账号缺少必要权限(如aiplatform.user、storage.objectViewer等)。
  • 依赖版本兼容问题:使用的google-cloud-aiplatform==1.25.0与langchain==0.0.187可能存在API不兼容,导致调用失败。

部署方式检查

  • Dockerfile:基础镜像选择合理,但安装了部分不必要的依赖(如unstructured、pdf2image),增加了镜像体积;未配置服务账号相关环境变量。
  • Cloud Build流程:镜像构建与推送步骤正确,但未加入镜像本地验证环节。
  • 端点部署:若未指定具备足够权限的服务账号,会导致容器内服务调用失败。

修复建议

1. 修正请求格式解析

修改FastAPI的predict接口,正确解析Vertex AI的请求结构:

async def predict(request: Request):
    body = await request.json()
    # 从instances数组中提取question字段
    question = body["instances"][0]["question"]
    # 后续逻辑保持不变

2. 优化资源初始化

将MatchingEngine和TextGenerationModel的初始化移到全局范围,避免重复创建:

# 全局初始化,仅执行一次
embeddings = VertexAIEmbeddings()
vector_store = MatchingEngine.from_components(
                    index_id=INDEX_ID,
                    region=REGION,
                    embedding=embeddings,
                    project_id=PROJECT_ID,
                    endpoint_id=ENDPOINT_ID,
                    gcs_bucket_name=DOCS_BUCKET)
model = TextGenerationModel.from_pretrained(TEXT_GENERATION_MODEL)

def matching_engine_search(question):
    relevant_documentation=vector_store.similarity_search(question, k=8)
    context = "\n".join([doc.page_content for doc in relevant_documentation])[:10000]
    return str(context)

3. 配置服务账号权限

确保部署端点使用的服务账号具备以下角色:

  • roles/aiplatform.user:允许访问Vertex AI服务
  • roles/storage.objectViewer:允许访问GCS桶内的文档
  • roles/aiplatform.indexEndpointViewer:允许访问Matching Engine端点

4. 调整依赖版本

升级google-cloud-aiplatform至较新版本(如>=1.30.0),并移除不必要的依赖,简化Dockerfile:

FROM tiangolo/uvicorn-gunicorn-fastapi:python3.8-slim
RUN pip install --no-cache-dir google-cloud-aiplatform>=1.30.0 langchain>=0.0.200 numpy>=1.24.0 pydantic>=1.10.0
COPY main.py ./main.py

5. 查看端点日志

通过GCP控制台进入Vertex AI端点页面,查看容器的日志详情,获取具体错误堆栈,精准定位问题。


附用户提供的代码文件

Python Code

import uvicorn

import os
import numpy as np

from fastapi import Request, FastAPI, Response
from fastapi.responses import JSONResponse

from langchain.vectorstores.matching_engine import MatchingEngine
from langchain.agents import Tool
from langchain.embeddings import VertexAIEmbeddings
from vertexai.preview.language_models import TextGenerationModel

embeddings = VertexAIEmbeddings()

INDEX_ID = "<index id>"
ENDPOINT_ID = "<index endpoint id>"
PROJECT_ID = '<project name>'
REGION = 'us-central1'
DOCS_BUCKET='<bucket name>'
TEXT_GENERATION_MODEL='text-bison@001'

def matching_engine_search(question):

    vector_store = MatchingEngine.from_components(
                        index_id=INDEX_ID,
                        region=REGION,
                        embedding=embeddings,
                        project_id=PROJECT_ID,
                        endpoint_id=ENDPOINT_ID,
                        gcs_bucket_name=DOCS_BUCKET)

    relevant_documentation=vector_store.similarity_search(question, k=8)
    context = "\n".join([doc.page_content for doc in relevant_documentation])[:10000]
    return str(context)

app = FastAPI(title="Chatbot")

AIP_HEALTH_ROUTE = os.environ.get('AIP_HEALTH_ROUTE', '/health')
AIP_PREDICT_ROUTE = os.environ.get('AIP_PREDICT_ROUTE', '/predict')

@app.get(AIP_HEALTH_ROUTE, status_code=200)
async def health():
    return {'health': 'ok'}

@app.post(AIP_PREDICT_ROUTE)
async def predict(request: Request):
    body = await request.json()
    print(body)

    question = body["question"]

    matching_engine_response=matching_engine_search(question)

    prompt=f"""
    Follow exactly those 3 steps:
    1. Read the context below and aggregrate this data
    Context : {matching_engine_response}
    2. Answer the question using only this context
    3. Show the source for your answers
    User Question: {question}


    If you don't have any context and are unsure of the answer, reply that you don't know about this topic.
    """

    model = TextGenerationModel.from_pretrained(TEXT_GENERATION_MODEL)
    response = model.predict(
            prompt,
            temperature=0.2,
            top_k=40,
            top_p=.8,
            max_output_tokens=1024,
    )

    print(f"Question: \n{question}")
    print(f"Response: \n{response.text}")

    return {"predictions": [{"response": response.text}] }

if __name__ == "__main__":
  uvicorn.run(app, host="0.0.0.0",port=8080)

Dockerfile

FROM tiangolo/uvicorn-gunicorn-fastapi:python3.8-slim
RUN pip install --no-cache-dir google-cloud-aiplatform==1.25.0 langchain==0.0.187 xmltodict==0.13.0 unstructured==0.7.0 pdf2image==1.16.3 numpy==1.23.1 pydantic==1.10.8 typing-inspect==0.8.0 typing_extensions==4.5.0
COPY main.py ./main.py

Cloudbuild.yaml

steps:
# Build the container image
- name: 'gcr.io/cloud-builders/docker'
  args: ['build', '-t', 'gcr.io/<project name>/chatbot', '.']
# Push the container image to Container Registry
- name: 'gcr.io/cloud-builders/docker'
  args: ['push', 'gcr.io/<project name>/chatbot']

images:
- gcr.io/<project name>/chatbot

查询端点代码

from google.cloud import aiplatform

aiplatform.init(project=PROJECT_ID,
                location=REGION)

instances = [{"question": "<Some question>"}]

endpoint = aiplatform.Endpoint("projects/<project id>/locations/us-central1/endpoints/<model endpoint id>")

prediction = endpoint.predict(instances=instances)
print(prediction)

内容的提问来源于stack exchange,提问作者user1758952

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 12:07:19