You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

部署大语言模型t0pp到SageMaker端点遇健康检查失败求助

部署BigScience T0pp到SageMaker端点失败:健康检查不通过的原因分析

用户尝试将BigScience T0pp大语言模型部署到SageMaker端点,使用的代码如下:

from sagemaker.huggingface import HuggingFaceModel
import sagemaker

role = sagemaker.get_execution_role()

hub = {
  'HF_MODEL_ID':'bigscience/T0', # model_id from hf.co/models
  'HF_TASK':'text2text-generation' # NLP task you want to use for predictions
}

# create Hugging Face Model Class
huggingface_model = HuggingFaceModel(
   env=hub,
   role=role, # iam role with permissions to create an Endpoint
   transformers_version="4.6", # transformers version used
   pytorch_version="1.7", # pytorch version used
   py_version="py36", # python version of the DLC
)

# deploy model to SageMaker Inference
predictor = huggingface_model.deploy(
   initial_instance_count=1,
   instance_type="ml.m5.xlarge"
)

部署时出现如下错误:

UnexpectedStatusException: Error hosting endpoint huggingface-pytorch-inference-2022-09-21-15-44-30-116: Failed. Reason: The primary container for production variant AllTraffic did not pass the ping health check.

可能的原因如下:

  • 资源规格不足:ml.m5.xlarge是CPU实例,T0pp属于参数规模较大的模型,CPU资源无法在健康检查超时内完成模型加载,导致容器启动失败。建议更换为GPU实例,比如ml.g4dn.xlarge、ml.p3.2xlarge等。
  • 依赖版本过旧:代码中指定的transformers 4.6、PyTorch 1.7版本过低,无法兼容T0pp模型的加载需求。建议升级transformers到4.10及以上版本,PyTorch到1.9及以上版本,同时将py_version调整为py38。
  • 模型ID错误:代码中填写的HF_MODEL_ID为bigscience/T0,但实际要部署的是T0pp,正确的模型ID应为bigscience/T0pp,错误的模型ID会导致加载异常。
  • 健康检查超时设置过短:大模型加载耗时较长,默认的健康检查超时时间不足以等待模型完全加载完成。可以通过自定义推理脚本,或者在部署时调整端点的健康检查配置延长超时时间。

内容的提问来源于stack exchange,提问作者Amir Imani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 10:40:34