You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

部署量化LLM至AWS SageMaker Endpoint时健康检查失败求助

部署量化LLM到AWS SageMaker端点时的健康检查失败问题

我将微调后的LLM模型(HuggingFaceH4/starchat-beta)的量化版本部署到AWS SageMaker Endpoint时,持续出现以下错误:

生产变体AllTraffic的主容器未通过ping健康检查,请查看该端点的CloudWatch日志。

查看CloudWatch日志后,发现关键报错:

java.io.FileNotFoundException: .py file not found in: /opt/ml/model

以下是我的部署代码,请求技术帮助:

!mkdir code
%%writefile code/inference.py
from typing import Dict, List, Any
import torch


def model_fn(model_dir):
    # load model and processor from model_dir
    # Activate 4-bit precision base model loading
    use_4bit = True
    # Compute dtype for 4-bit base models
    bnb_4bit_compute_dtype = "float16"
    # Quantization type (fp4 or nf4)
    bnb_4bit_quant_type = "nf4"
    # Activate nested quantization for 4-bit base models (double quantization)
    use_nested_quant = True
    # Model Name
    # model_name = "MODEL_HF_PATH"

    # Load tokenizer and model with QLoRA configuration
    compute_dtype = getattr(torch, bnb_4bit_compute_dtype)

    bnb_config = BitsAndBytesConfig(
       load_in_4bit=True,
       bnb_4bit_quant_type="nf4",
       bnb_4bit_use_double_quant=True,
       bnb_4bit_compute_dtype=torch.bfloat16
    )

    # Load base model
    llm_model = AutoModelForCausalLM.from_pretrained(
        model_dir,
        quantization_config=bnb_config,
        device_map="auto",

    )

    tokenizer = AutoTokenizer.from_pretrained(model_dir))

    return llm_model, tokenizer

from distutils.dir_util import copy_tree
from pathlib import Path
from tempfile import TemporaryDirectory
from huggingface_hub import snapshot_download

HF_MODEL_ID='MODEL_HF_PATH'
# create model dir
model_tar_dir = Path(HF_MODEL_ID.split("/")[-1])
model_tar_dir.mkdir()

# setup temporary directory
with TemporaryDirectory() as tmpdir:
    # download snapshot
    snapshot_dir = snapshot_download(repo_id=HF_MODEL_ID, cache_dir=tmpdir,resume_download=True)
    # copy snapshot to model dir
    print('Copying...')
    copy_tree(snapshot_dir, str(model_tar_dir))

copy_tree("code/", str(model_tar_dir.joinpath("code")))

import tarfile
import os

# helper to create the model.tar.gz
def compress(tar_dir=None,output_file="model.tar.gz"):
    parent_dir=os.getcwd()
    os.chdir(tar_dir)
    with tarfile.open(os.path.join(parent_dir, output_file), "w:gz") as tar:
        for item in os.listdir('.'):
          print(item)
          tar.add(item, arcname=item)
    os.chdir(parent_dir)

compress(str(model_tar_dir))

from sagemaker.s3 import S3Uploader
# upload model.tar.gz to s3
s3_model_uri = S3Uploader.upload(local_path="model.tar.gz", desired_s3_uri=f"s3://{sess.default_bucket()}/model_name")

from sagemaker.huggingface.model import HuggingFaceModel


# create Hugging Face Model Class
huggingface_model = HuggingFaceModel(
    model_data=s3_model_uri,      # path to your model and script
    role=role, # iam role with permissions to create an Endpoint
    image_uri = '763104351884.dkr.ecr.us-east-2.amazonaws.com/djl-inference:0.23.0-fastertransformer5.3.0-cu118'
)

# deploy the endpoint endpoint
predictor = huggingface_model.deploy(
    initial_instance_count=1,
    instance_type="ml.g5.xlarge"
    )

问题分析与修复方案

核心原因

你使用的是DJL推理镜像(djl-inference),但打包逻辑是按照Hugging Face原生推理脚本的方式设计的,两者对推理脚本的位置要求不匹配:

  • DJL要求推理脚本直接放在模型根目录(/opt/ml/model)下,而你把inference.py放在了code/子目录中,导致镜像找不到脚本文件。

修复方案

方案1:调整脚本位置适配DJL镜像

  1. 修改打包逻辑,将inference.py直接复制到模型根目录,替换原有的copy_tree("code/", ...)代码:
    # 替换原copy_tree("code/", str(model_tar_dir.joinpath("code")))
    import shutil
    shutil.copy("code/inference.py", str(model_tar_dir))
    
  2. 确保model.tar.gz解压后的结构如下:
    model.tar.gz/
    ├── inference.py
    ├── config.json
    ├── pytorch_model.bin(或量化后的模型文件)
    ├── tokenizer.json
    └── ...(其他模型相关文件)
    

方案2:更换为Hugging Face官方推理镜像

如果想继续保留code/目录的结构,改用Hugging Face提供的SageMaker推理镜像,修改HuggingFaceModel的image_uri(以us-east-2区域为例):

image_uri = '763104351884.dkr.ecr.us-east-2.amazonaws.com/huggingface-pytorch-inference:2.0.0-transformers4.28.1-gpu-py310-cu118-ubuntu20.04'

额外需要修正的问题

  1. inference.py存在语法错误:tokenizer = AutoTokenizer.from_pretrained(model_dir)) 多了一个右括号,修正为:
    tokenizer = AutoTokenizer.from_pretrained(model_dir)
    
  2. 缺少必要导入,需在inference.py开头添加:
    from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
    

内容的提问来源于stack exchange,提问作者RalphCh97

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 03:42:39