You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

微调后Falcon-7B部署AWS SageMaker端点报错求助

Falcon-7B微调后无法部署到AWS SageMaker端点:内存不足及模型类型不兼容问题

问题详情

在AWS SageMaker上完成Falcon-7B模型的微调训练后,未压缩的权重已成功存储到S3桶,但部署至SageMaker端点时失败。核心问题是部署容器无法通过健康检查,CloudWatch日志显示内存不足以处理预填充tokens,即使降低token数量仍报错;尝试将模型推至Hugging Face再部署时,又出现模型类型不兼容错误。

复现步骤

  1. 获取LLM镜像URI
from sagemaker.huggingface import get_huggingface_llm_image_uri

# retrieve the llm image uri
llm_image = get_huggingface_llm_image_uri(
  "huggingface",
  version="1.1.0",
  session=sess,
)

# print ecr image uri
print(f"llm image uri: {llm_image}")
  1. 创建HuggingFaceModel并指定模型路径
import json
from sagemaker.huggingface import HuggingFaceModel

model_s3_path = huggingface_estimator.model_data["S3DataSource"]["S3Uri"]

# sagemaker config
instance_type = "ml.g5.12xlarge"
number_of_gpu = 1
health_check_timeout = 300

# Define Model and Endpoint configuration parameter
config = {
  'HF_MODEL_ID': "/opt/ml/model", # path to where sagemaker stores the model
  'SM_NUM_GPUS': json.dumps(number_of_gpu), # Number of GPU used per replica
  'MAX_INPUT_LENGTH': json.dumps(1024), # Max length of input text
  'MAX_TOTAL_TOKENS': json.dumps(2048), # Max length of the generation (including input text)
}

# create HuggingFaceModel with the image uri
llm_model = HuggingFaceModel(
  role=role,
  image_uri=llm_image,
  model_data={'S3DataSource':{'S3Uri': model_s3_path,'S3DataType': 'S3Prefix','CompressionType': 'None'}},
  env=config
)
  1. 部署模型
llm = llm_model.deploy(
  initial_instance_count=1,
  instance_type=instance_type,
  container_startup_health_check_timeout=health_check_timeout, # 10 minutes to be able to load the model
)

错误信息

  1. 部署时的端点错误:
UnexpectedStatusException: Error hosting endpoint huggingface-pytorch-tgi-inference-2023-10-21-16-47-53-072: Failed. Reason: The primary container for production variant AllTraffic did not pass the ping health check. Please check CloudWatch logs for this endpoint..
  1. CloudWatch核心错误:
RuntimeError: Not enough memory to handle 4096 prefill tokens. You need to decrease `--max-batch-prefill-tokens`

将预填充token减半至2048后仍报错:

RuntimeError: Not enough memory to handle 2048 prefill tokens. You need to decrease `--max-batch-prefill-tokens`
  1. 推至Hugging Face后部署的错误:
ValueError: Unsupported model type falcon

已尝试的解决方案

  • 降低MAX_INPUT_LENGTH和MAX_TOTAL_TOKENS参数值
  • 将训练后的模型推至Hugging Face Hub后再部署
  • 将模型打包为model.tar.gz压缩包形式部署

解决方案

1. 针对内存不足问题:调整TGI内存相关环境变量

Falcon-7B在SageMaker TGI容器中加载时,仅调整输入token上限不足以解决内存问题,需添加量化及批量token限制参数:

config = {
  'HF_MODEL_ID': "/opt/ml/model",
  'SM_NUM_GPUS': json.dumps(number_of_gpu),
  'MAX_INPUT_LENGTH': json.dumps(1024),
  'MAX_TOTAL_TOKENS': json.dumps(2048),
  'MAX_BATCH_PREFILL_TOKENS': json.dumps(1024),  # 显式设置预填充token上限
  'MAX_BATCH_TOTAL_TOKENS': json.dumps(2048),    # 限制单批次总token数
  'LOAD_IN_4BIT': "true",                        # 启用4bit量化大幅降低显存占用
}

4bit量化可将Falcon-7B的显存占用从FP16的约13GB降至约4GB,直接缓解内存压力。

2. 针对模型类型不兼容问题:添加远程代码信任配置

Falcon模型依赖自定义代码实现,推至Hugging Face Hub后部署时,必须启用信任远程代码的配置:

config = {
  'HF_MODEL_ID': "your-hf-username/your-falcon-model",
  'SM_NUM_GPUS': json.dumps(number_of_gpu),
  'MAX_INPUT_LENGTH': json.dumps(1024),
  'MAX_TOTAL_TOKENS': json.dumps(2048),
  'HF_TRUST_REMOTE_CODE': "true",  # 必须添加,加载Falcon自定义代码
  'LOAD_IN_4BIT': "true",
}

3. 模型打包优化:确保压缩包结构正确

使用model.tar.gz部署时,需确保压缩包解压后直接包含模型核心文件(pytorch_model.bin、config.json等),而非嵌套子目录。打包命令示例:

cd /path/to/model-files
tar -czf model.tar.gz *

上传至S3后,部署时需指定该压缩包的S3路径,而非S3前缀。

4. 实例资源优化:启用张量并行

如果使用多GPU实例,可启用张量并行进一步分摊内存压力:

config = {
  'HF_MODEL_ID': "/opt/ml/model",
  'SM_NUM_GPUS': json.dumps(2),
  'MAX_INPUT_LENGTH': json.dumps(1024),
  'MAX_TOTAL_TOKENS': json.dumps(2048),
  'MAX_BATCH_PREFILL_TOKENS': json.dumps(1024),
  'LOAD_IN_4BIT': "true",
  'TP_SIZE': json.dumps(2),  # 张量并行数与GPU数量一致
}

内容的提问来源于stack exchange,提问作者Jacob Brophy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 05:54:54