Vertex AI端点部署自定义容器后查询报500内部服务器错误排查
Vertex AI自定义容器端点500内部服务器错误排查
问题描述
已成功将搭载PaLM大语言模型的自定义容器部署到Vertex AI端点,但通过Vertex AI API或gcloud cli查询端点时,收到500 Internal Server Error回复。
错误信息(翻译后)
500 Internal Server Error
无法获取预测结果。请检查端点日志以获取详细信息。
可能的原因分析
- 请求格式解析错误:Vertex AI端点发送的预测请求结构为
{"instances": [{"question": "..."}]},但当前代码直接读取body["question"],会因找不到对应字段引发KeyError,导致500错误。 - 资源重复初始化:每次预测请求都重新创建
MatchingEngine和TextGenerationModel实例,会增加资源开销,甚至引发初始化超时。 - 服务账号权限不足:容器内代码调用Vertex AI的Matching Engine、文本生成模型及GCS桶时,部署端点使用的服务账号缺少必要权限(如
aiplatform.user、storage.objectViewer等)。 - 依赖版本兼容问题:使用的
google-cloud-aiplatform==1.25.0与langchain==0.0.187可能存在API不兼容,导致调用失败。
部署方式检查
- Dockerfile:基础镜像选择合理,但安装了部分不必要的依赖(如
unstructured、pdf2image),增加了镜像体积;未配置服务账号相关环境变量。 - Cloud Build流程:镜像构建与推送步骤正确,但未加入镜像本地验证环节。
- 端点部署:若未指定具备足够权限的服务账号,会导致容器内服务调用失败。
修复建议
1. 修正请求格式解析
修改FastAPI的predict接口,正确解析Vertex AI的请求结构:
async def predict(request: Request): body = await request.json() # 从instances数组中提取question字段 question = body["instances"][0]["question"] # 后续逻辑保持不变
2. 优化资源初始化
将MatchingEngine和TextGenerationModel的初始化移到全局范围,避免重复创建:
# 全局初始化,仅执行一次 embeddings = VertexAIEmbeddings() vector_store = MatchingEngine.from_components( index_id=INDEX_ID, region=REGION, embedding=embeddings, project_id=PROJECT_ID, endpoint_id=ENDPOINT_ID, gcs_bucket_name=DOCS_BUCKET) model = TextGenerationModel.from_pretrained(TEXT_GENERATION_MODEL) def matching_engine_search(question): relevant_documentation=vector_store.similarity_search(question, k=8) context = "\n".join([doc.page_content for doc in relevant_documentation])[:10000] return str(context)
3. 配置服务账号权限
确保部署端点使用的服务账号具备以下角色:
roles/aiplatform.user:允许访问Vertex AI服务roles/storage.objectViewer:允许访问GCS桶内的文档roles/aiplatform.indexEndpointViewer:允许访问Matching Engine端点
4. 调整依赖版本
升级google-cloud-aiplatform至较新版本(如>=1.30.0),并移除不必要的依赖,简化Dockerfile:
FROM tiangolo/uvicorn-gunicorn-fastapi:python3.8-slim RUN pip install --no-cache-dir google-cloud-aiplatform>=1.30.0 langchain>=0.0.200 numpy>=1.24.0 pydantic>=1.10.0 COPY main.py ./main.py
5. 查看端点日志
通过GCP控制台进入Vertex AI端点页面,查看容器的日志详情,获取具体错误堆栈,精准定位问题。
附用户提供的代码文件
Python Code
import uvicorn import os import numpy as np from fastapi import Request, FastAPI, Response from fastapi.responses import JSONResponse from langchain.vectorstores.matching_engine import MatchingEngine from langchain.agents import Tool from langchain.embeddings import VertexAIEmbeddings from vertexai.preview.language_models import TextGenerationModel embeddings = VertexAIEmbeddings() INDEX_ID = "<index id>" ENDPOINT_ID = "<index endpoint id>" PROJECT_ID = '<project name>' REGION = 'us-central1' DOCS_BUCKET='<bucket name>' TEXT_GENERATION_MODEL='text-bison@001' def matching_engine_search(question): vector_store = MatchingEngine.from_components( index_id=INDEX_ID, region=REGION, embedding=embeddings, project_id=PROJECT_ID, endpoint_id=ENDPOINT_ID, gcs_bucket_name=DOCS_BUCKET) relevant_documentation=vector_store.similarity_search(question, k=8) context = "\n".join([doc.page_content for doc in relevant_documentation])[:10000] return str(context) app = FastAPI(title="Chatbot") AIP_HEALTH_ROUTE = os.environ.get('AIP_HEALTH_ROUTE', '/health') AIP_PREDICT_ROUTE = os.environ.get('AIP_PREDICT_ROUTE', '/predict') @app.get(AIP_HEALTH_ROUTE, status_code=200) async def health(): return {'health': 'ok'} @app.post(AIP_PREDICT_ROUTE) async def predict(request: Request): body = await request.json() print(body) question = body["question"] matching_engine_response=matching_engine_search(question) prompt=f""" Follow exactly those 3 steps: 1. Read the context below and aggregrate this data Context : {matching_engine_response} 2. Answer the question using only this context 3. Show the source for your answers User Question: {question} If you don't have any context and are unsure of the answer, reply that you don't know about this topic. """ model = TextGenerationModel.from_pretrained(TEXT_GENERATION_MODEL) response = model.predict( prompt, temperature=0.2, top_k=40, top_p=.8, max_output_tokens=1024, ) print(f"Question: \n{question}") print(f"Response: \n{response.text}") return {"predictions": [{"response": response.text}] } if __name__ == "__main__": uvicorn.run(app, host="0.0.0.0",port=8080)
Dockerfile
FROM tiangolo/uvicorn-gunicorn-fastapi:python3.8-slim RUN pip install --no-cache-dir google-cloud-aiplatform==1.25.0 langchain==0.0.187 xmltodict==0.13.0 unstructured==0.7.0 pdf2image==1.16.3 numpy==1.23.1 pydantic==1.10.8 typing-inspect==0.8.0 typing_extensions==4.5.0 COPY main.py ./main.py
Cloudbuild.yaml
steps: # Build the container image - name: 'gcr.io/cloud-builders/docker' args: ['build', '-t', 'gcr.io/<project name>/chatbot', '.'] # Push the container image to Container Registry - name: 'gcr.io/cloud-builders/docker' args: ['push', 'gcr.io/<project name>/chatbot'] images: - gcr.io/<project name>/chatbot
查询端点代码
from google.cloud import aiplatform aiplatform.init(project=PROJECT_ID, location=REGION) instances = [{"question": "<Some question>"}] endpoint = aiplatform.Endpoint("projects/<project id>/locations/us-central1/endpoints/<model endpoint id>") prediction = endpoint.predict(instances=instances) print(prediction)
内容的提问来源于stack exchange,提问作者user1758952
相关产品推荐
相关产品推荐

