部署大语言模型t0pp到SageMaker端点遇健康检查失败求助
部署BigScience T0pp到SageMaker端点失败:健康检查不通过的原因分析
用户尝试将BigScience T0pp大语言模型部署到SageMaker端点,使用的代码如下:
from sagemaker.huggingface import HuggingFaceModel import sagemaker role = sagemaker.get_execution_role() hub = { 'HF_MODEL_ID':'bigscience/T0', # model_id from hf.co/models 'HF_TASK':'text2text-generation' # NLP task you want to use for predictions } # create Hugging Face Model Class huggingface_model = HuggingFaceModel( env=hub, role=role, # iam role with permissions to create an Endpoint transformers_version="4.6", # transformers version used pytorch_version="1.7", # pytorch version used py_version="py36", # python version of the DLC ) # deploy model to SageMaker Inference predictor = huggingface_model.deploy( initial_instance_count=1, instance_type="ml.m5.xlarge" )
部署时出现如下错误:
UnexpectedStatusException: Error hosting endpoint huggingface-pytorch-inference-2022-09-21-15-44-30-116: Failed. Reason: The primary container for production variant AllTraffic did not pass the ping health check.
可能的原因如下:
- 资源规格不足:ml.m5.xlarge是CPU实例,T0pp属于参数规模较大的模型,CPU资源无法在健康检查超时内完成模型加载,导致容器启动失败。建议更换为GPU实例,比如ml.g4dn.xlarge、ml.p3.2xlarge等。
- 依赖版本过旧:代码中指定的transformers 4.6、PyTorch 1.7版本过低,无法兼容T0pp模型的加载需求。建议升级transformers到4.10及以上版本,PyTorch到1.9及以上版本,同时将py_version调整为py38。
- 模型ID错误:代码中填写的
HF_MODEL_ID为bigscience/T0,但实际要部署的是T0pp,正确的模型ID应为bigscience/T0pp,错误的模型ID会导致加载异常。 - 健康检查超时设置过短:大模型加载耗时较长,默认的健康检查超时时间不足以等待模型完全加载完成。可以通过自定义推理脚本,或者在部署时调整端点的健康检查配置延长超时时间。
内容的提问来源于stack exchange,提问作者Amir Imani
相关产品推荐
相关产品推荐

