无需HuggingFace Token部署SageMaker Llama2端点失败求助
问题解决:无需HuggingFace部署Llama-2原生模型到SageMaker端点
问题根源
你使用的HuggingFace LMI镜像默认适配HuggingFace格式模型(即config.json、pytorch_model.bin这类结构),但你的模型是Meta原生的.pth格式,缺少必要的推理适配脚本和配置,导致部署流程中断,端点无法创建。
解决方案步骤
1. 调整模型归档包结构
在你的model.tar.gz中添加code/目录,包含推理必需的脚本和配置:
code/inference.py:负责加载模型、处理请求code/requirements.txt:声明依赖(如torch、sentencepiece等)- (若使用LMI镜像)
serving.properties:配置原生模型加载规则
示例inference.py(原生PyTorch加载Llama-2)
import torch import sentencepiece as spm from pathlib import Path model_dir = "/opt/ml/model" # 加载tokenizer tokenizer = spm.SentencePieceProcessor(model_file=str(Path(model_dir)/"tokenizer.model")) # 加载Llama-2原生模型(需将Meta官方Llama仓库的llama目录打包进code/) from llama.model import Transformer params = torch.load(Path(model_dir)/"params.json") model = Transformer(**params) model.load_state_dict(torch.load(Path(model_dir)/"consolidated.00.pth"), strict=False) model.eval().to("cuda") def model_fn(model_dir): return {"model": model, "tokenizer": tokenizer} def predict_fn(input_data, model): inputs = input_data.get("inputs") tokenized = model["tokenizer"].encode(inputs, add_bos=True, return_tensors="pt").to("cuda") with torch.no_grad(): outputs = model["model"].generate(tokenized, max_new_tokens=200) response = model["tokenizer"].decode(outputs[0], skip_special_tokens=True) return {"outputs": response}
示例requirements.txt
torch>=2.0.0 sentencepiece>=0.1.97
2. 选择适配的镜像
- 纯PyTorch部署:使用SageMaker PyTorch DLC镜像,示例(根据区域调整):
763104351884.dkr.ecr.us-east-1.amazonaws.com/pytorch-inference:2.0.0-gpu-py310 - LMI镜像适配:添加
serving.properties配置原生模型加载:engine=PyTorch option.model_type=llama option.tensor_parallel_degree=1 option.load_format=pth
3. 纯boto3部署流程
import boto3 import time sagemaker = boto3.client("sagemaker") region = boto3.Session().region_name # 1. 创建模型 model_name = "llama2-native-model" role_arn = "你的SageMaker执行角色ARN" model_data_url = "s3://你的存储桶路径/model.tar.gz" image_uri = f"763104351884.dkr.ecr.{region}.amazonaws.com/pytorch-inference:2.0.0-gpu-py310" sagemaker.create_model( ModelName=model_name, ExecutionRoleArn=role_arn, PrimaryContainer={ "Image": image_uri, "ModelDataUrl": model_data_url, "Environment": { "SAGEMAKER_PROGRAM": "inference.py", "SAGEMAKER_SUBMIT_DIRECTORY": "/opt/ml/model/code" } } ) # 2. 创建端点配置 endpoint_config_name = "llama2-endpoint-config" sagemaker.create_endpoint_config( EndpointConfigName=endpoint_config_name, ProductionVariants=[ { "VariantName": "primary", "ModelName": model_name, "InstanceType": "ml.g5.2xlarge", "InitialInstanceCount": 1 } ] ) # 3. 创建端点并等待就绪 endpoint_name = "ss-llama2-endpoint" sagemaker.create_endpoint( EndpointName=endpoint_name, EndpointConfigName=endpoint_config_name ) while True: status = sagemaker.describe_endpoint(EndpointName=endpoint_name)["EndpointStatus"] print(f"端点状态: {status}") if status == "InService": break time.sleep(30)
4. 测试端点
runtime = boto3.client("sagemaker-runtime") response = runtime.invoke_endpoint( EndpointName=endpoint_name, ContentType="application/json", Body='{"inputs": "What is AWS SageMaker?"}' ) result = response["Body"].read().decode("utf-8") print(result)
关键注意事项
- 确保
code/目录结构正确,SageMaker会自动解压并执行指定的推理脚本 - 需将Meta官方Llama仓库的
llama/目录打包进code/,才能加载原生模型 - 部署失败时,可通过SageMaker控制台查看端点日志,排查依赖缺失、模型加载失败等具体错误
内容的提问来源于stack exchange,提问作者Spas Kalaydzhiyski
相关产品推荐
相关产品推荐

