You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无需HuggingFace Token部署SageMaker Llama2端点失败求助

问题解决:无需HuggingFace部署Llama-2原生模型到SageMaker端点

问题根源

你使用的HuggingFace LMI镜像默认适配HuggingFace格式模型(即config.json、pytorch_model.bin这类结构),但你的模型是Meta原生的.pth格式,缺少必要的推理适配脚本和配置,导致部署流程中断,端点无法创建。

解决方案步骤

1. 调整模型归档包结构

在你的model.tar.gz中添加code/目录,包含推理必需的脚本和配置:

  • code/inference.py:负责加载模型、处理请求
  • code/requirements.txt:声明依赖(如torch、sentencepiece等)
  • (若使用LMI镜像)serving.properties:配置原生模型加载规则

示例inference.py(原生PyTorch加载Llama-2)

import torch
import sentencepiece as spm
from pathlib import Path

model_dir = "/opt/ml/model"

# 加载tokenizer
tokenizer = spm.SentencePieceProcessor(model_file=str(Path(model_dir)/"tokenizer.model"))

# 加载Llama-2原生模型(需将Meta官方Llama仓库的llama目录打包进code/)
from llama.model import Transformer
params = torch.load(Path(model_dir)/"params.json")
model = Transformer(**params)
model.load_state_dict(torch.load(Path(model_dir)/"consolidated.00.pth"), strict=False)
model.eval().to("cuda")

def model_fn(model_dir):
    return {"model": model, "tokenizer": tokenizer}

def predict_fn(input_data, model):
    inputs = input_data.get("inputs")
    tokenized = model["tokenizer"].encode(inputs, add_bos=True, return_tensors="pt").to("cuda")
    with torch.no_grad():
        outputs = model["model"].generate(tokenized, max_new_tokens=200)
    response = model["tokenizer"].decode(outputs[0], skip_special_tokens=True)
    return {"outputs": response}

示例requirements.txt

torch>=2.0.0
sentencepiece>=0.1.97

2. 选择适配的镜像

  • 纯PyTorch部署:使用SageMaker PyTorch DLC镜像,示例(根据区域调整):
    763104351884.dkr.ecr.us-east-1.amazonaws.com/pytorch-inference:2.0.0-gpu-py310
  • LMI镜像适配:添加serving.properties配置原生模型加载:
    engine=PyTorch
    option.model_type=llama
    option.tensor_parallel_degree=1
    option.load_format=pth
    

3. 纯boto3部署流程

import boto3
import time

sagemaker = boto3.client("sagemaker")
region = boto3.Session().region_name

# 1. 创建模型
model_name = "llama2-native-model"
role_arn = "你的SageMaker执行角色ARN"
model_data_url = "s3://你的存储桶路径/model.tar.gz"
image_uri = f"763104351884.dkr.ecr.{region}.amazonaws.com/pytorch-inference:2.0.0-gpu-py310"

sagemaker.create_model(
    ModelName=model_name,
    ExecutionRoleArn=role_arn,
    PrimaryContainer={
        "Image": image_uri,
        "ModelDataUrl": model_data_url,
        "Environment": {
            "SAGEMAKER_PROGRAM": "inference.py",
            "SAGEMAKER_SUBMIT_DIRECTORY": "/opt/ml/model/code"
        }
    }
)

# 2. 创建端点配置
endpoint_config_name = "llama2-endpoint-config"
sagemaker.create_endpoint_config(
    EndpointConfigName=endpoint_config_name,
    ProductionVariants=[
        {
            "VariantName": "primary",
            "ModelName": model_name,
            "InstanceType": "ml.g5.2xlarge",
            "InitialInstanceCount": 1
        }
    ]
)

# 3. 创建端点并等待就绪
endpoint_name = "ss-llama2-endpoint"
sagemaker.create_endpoint(
    EndpointName=endpoint_name,
    EndpointConfigName=endpoint_config_name
)

while True:
    status = sagemaker.describe_endpoint(EndpointName=endpoint_name)["EndpointStatus"]
    print(f"端点状态: {status}")
    if status == "InService":
        break
    time.sleep(30)

4. 测试端点

runtime = boto3.client("sagemaker-runtime")

response = runtime.invoke_endpoint(
    EndpointName=endpoint_name,
    ContentType="application/json",
    Body='{"inputs": "What is AWS SageMaker?"}'
)

result = response["Body"].read().decode("utf-8")
print(result)

关键注意事项

  • 确保code/目录结构正确,SageMaker会自动解压并执行指定的推理脚本
  • 需将Meta官方Llama仓库的llama/目录打包进code/,才能加载原生模型
  • 部署失败时,可通过SageMaker控制台查看端点日志,排查依赖缺失、模型加载失败等具体错误

内容的提问来源于stack exchange,提问作者Spas Kalaydzhiyski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 12:51:07