You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在AWS SageMaker部署微调Gemma 7B模型遇阻,求部署方案及镜像URI

部署LoRA微调后的Gemma 7B到AWS SageMaker端点

适配的镜像URI

针对transformers 4.38.0版本,选择AWS官方的HuggingFace推理DLC镜像,根据实例类型和区域调整:

GPU实例(推荐,Gemma 7B需显存支持)

763104351884.dkr.ecr.<你的AWS区域>.amazonaws.com/huggingface-pytorch-inference:2.0.0-transformers4.38.0-gpu-py310-cu118-ubuntu20.04

示例:us-east-1区域的镜像URI为 763104351884.dkr.ecr.us-east-1.amazonaws.com/huggingface-pytorch-inference:2.0.0-transformers4.38.0-gpu-py310-cu118-ubuntu20.04

CPU实例(仅用于测试,速度极慢)

763104351884.dkr.ecr.<你的AWS区域>.amazonaws.com/huggingface-pytorch-inference:2.0.0-transformers4.38.0-cpu-py310-ubuntu20.04

部署前的模型包调整

  1. 清理冗余文件:从你的finetuned_gemma.tar.gz中删除code/.ipynb_checkpoints/目录,避免SageMaker加载时出现路径错误。
  2. 更新requirements.txt:确保code/requirements.txt包含依赖项:
transformers==4.38.0
accelerate>=0.27.0
bitsandbytes>=0.41.0
torch>=2.0.0
sentencepiece

如果你的LoRA权重未合并到基础模型,需额外添加peft>=0.8.0。

  1. 调整推理脚本(inference.py)
    根据你的模型是否合并了LoRA权重,选择对应脚本:

情况1:已合并LoRA权重到基础模型(你的文件结构显示为完整模型权重)

from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig
import torch

def model_fn(model_dir):
    # 加载Gemma分词器
    tokenizer = AutoTokenizer.from_pretrained(model_dir)
    # 加载微调后的模型,使用bf16节省显存
    model = AutoModelForCausalLM.from_pretrained(
        model_dir,
        torch_dtype=torch.bfloat16,
        device_map="auto",
        trust_remote_code=True
    )
    # 加载生成配置
    generation_config = GenerationConfig.from_pretrained(model_dir)
    return model, tokenizer, generation_config

def predict_fn(input_data, model_and_tokenizer):
    model, tokenizer, generation_config = model_and_tokenizer
    prompt = input_data["inputs"]
    # 适配Gemma的对话格式
    formatted_prompt = f"<start_of_turn>user\n{prompt}<end_of_turn>\n<start_of_turn>model\n"
    inputs = tokenizer(formatted_prompt, return_tensors="pt").to(model.device)
    
    with torch.no_grad():
        outputs = model.generate(
            **inputs,
            generation_config=generation_config,
            max_new_tokens=512,
            do_sample=True,
            temperature=0.7,
            top_p=0.9
        )
    
    response = tokenizer.decode(outputs[0], skip_special_tokens=True)
    # 提取模型输出内容
    model_response = response.split("<start_of_turn>model\n")[-1]
    return {"generated_text": model_response}

情况2:LoRA权重未合并(需单独加载基础模型)

from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig
from peft import PeftModel
import torch

def model_fn(model_dir):
    tokenizer = AutoTokenizer.from_pretrained(model_dir)
    # 加载官方Gemma 7B基础模型
    base_model = AutoModelForCausalLM.from_pretrained(
        "google/gemma-7b",
        torch_dtype=torch.bfloat16,
        device_map="auto",
        trust_remote_code=True
    )
    # 加载LoRA微调权重
    model = PeftModel.from_pretrained(base_model, model_dir)
    # 合并权重提升推理速度(可选)
    model = model.merge_and_unload()
    generation_config = GenerationConfig.from_pretrained(model_dir)
    return model, tokenizer, generation_config

# predict_fn与情况1完全一致

部署代码实现

使用SageMaker Python SDK完成部署:

import sagemaker
from sagemaker.huggingface import HuggingFaceModel
from sagemaker import get_execution_role

# 获取SageMaker执行角色
role = get_execution_role()

# 模型包在S3的存储路径
model_s3_path = "s3://你的存储桶名称/路径/finetuned_gemma.tar.gz"

# 替换为你的区域对应的镜像URI
image_uri = "763104351884.dkr.ecr.us-east-1.amazonaws.com/huggingface-pytorch-inference:2.0.0-transformers4.38.0-gpu-py310-cu118-ubuntu20.04"

# 创建HuggingFace模型对象
hf_model = HuggingFaceModel(
    model_data=model_s3_path,
    role=role,
    image_uri=image_uri,
    env={
        "HF_TASK": "text-generation",
        "SM_NUM_GPUS": "1"  # 对应实例的GPU数量,比如ml.g5.2xlarge为1
    }
)

# 部署到端点
predictor = hf_model.deploy(
    initial_instance_count=1,
    instance_type="ml.g5.2xlarge",  # 推荐实例,显存24GB足够运行Gemma 7B
    endpoint_name="gemma-7b-finetuned-lora-endpoint"
)

# 测试推理
test_input = {"inputs": "请解释大语言模型的LoRA微调原理"}
response = predictor.predict(test_input)
print(response["generated_text"])

关键注意事项

  • 实例选择:Gemma 7B需至少16GB显存,推荐使用ml.g5.2xlarge或更高配置的GPU实例。
  • 区域适配:镜像URI中的区域需与你的S3桶和SageMaker端点所在区域一致。
  • 权限配置:确保SageMaker角色拥有访问目标S3桶的权限,避免模型加载失败。
  • 量化选项:若显存不足,可在AutoModelForCausalLM.from_pretrained中添加load_in_4bit=True开启4bit量化,需确保bitsandbytes版本符合要求。

内容的提问来源于stack exchange,提问作者Sai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 23:54:53