微调后Falcon-7B部署AWS SageMaker端点报错求助
Falcon-7B微调后无法部署到AWS SageMaker端点:内存不足及模型类型不兼容问题
问题详情
在AWS SageMaker上完成Falcon-7B模型的微调训练后,未压缩的权重已成功存储到S3桶,但部署至SageMaker端点时失败。核心问题是部署容器无法通过健康检查,CloudWatch日志显示内存不足以处理预填充tokens,即使降低token数量仍报错;尝试将模型推至Hugging Face再部署时,又出现模型类型不兼容错误。
复现步骤
- 获取LLM镜像URI
from sagemaker.huggingface import get_huggingface_llm_image_uri # retrieve the llm image uri llm_image = get_huggingface_llm_image_uri( "huggingface", version="1.1.0", session=sess, ) # print ecr image uri print(f"llm image uri: {llm_image}")
- 创建HuggingFaceModel并指定模型路径
import json from sagemaker.huggingface import HuggingFaceModel model_s3_path = huggingface_estimator.model_data["S3DataSource"]["S3Uri"] # sagemaker config instance_type = "ml.g5.12xlarge" number_of_gpu = 1 health_check_timeout = 300 # Define Model and Endpoint configuration parameter config = { 'HF_MODEL_ID': "/opt/ml/model", # path to where sagemaker stores the model 'SM_NUM_GPUS': json.dumps(number_of_gpu), # Number of GPU used per replica 'MAX_INPUT_LENGTH': json.dumps(1024), # Max length of input text 'MAX_TOTAL_TOKENS': json.dumps(2048), # Max length of the generation (including input text) } # create HuggingFaceModel with the image uri llm_model = HuggingFaceModel( role=role, image_uri=llm_image, model_data={'S3DataSource':{'S3Uri': model_s3_path,'S3DataType': 'S3Prefix','CompressionType': 'None'}}, env=config )
- 部署模型
llm = llm_model.deploy( initial_instance_count=1, instance_type=instance_type, container_startup_health_check_timeout=health_check_timeout, # 10 minutes to be able to load the model )
错误信息
- 部署时的端点错误:
UnexpectedStatusException: Error hosting endpoint huggingface-pytorch-tgi-inference-2023-10-21-16-47-53-072: Failed. Reason: The primary container for production variant AllTraffic did not pass the ping health check. Please check CloudWatch logs for this endpoint..
- CloudWatch核心错误:
RuntimeError: Not enough memory to handle 4096 prefill tokens. You need to decrease `--max-batch-prefill-tokens`
将预填充token减半至2048后仍报错:
RuntimeError: Not enough memory to handle 2048 prefill tokens. You need to decrease `--max-batch-prefill-tokens`
- 推至Hugging Face后部署的错误:
ValueError: Unsupported model type falcon
已尝试的解决方案
- 降低
MAX_INPUT_LENGTH和MAX_TOTAL_TOKENS参数值 - 将训练后的模型推至Hugging Face Hub后再部署
- 将模型打包为
model.tar.gz压缩包形式部署
解决方案
1. 针对内存不足问题:调整TGI内存相关环境变量
Falcon-7B在SageMaker TGI容器中加载时,仅调整输入token上限不足以解决内存问题,需添加量化及批量token限制参数:
config = { 'HF_MODEL_ID': "/opt/ml/model", 'SM_NUM_GPUS': json.dumps(number_of_gpu), 'MAX_INPUT_LENGTH': json.dumps(1024), 'MAX_TOTAL_TOKENS': json.dumps(2048), 'MAX_BATCH_PREFILL_TOKENS': json.dumps(1024), # 显式设置预填充token上限 'MAX_BATCH_TOTAL_TOKENS': json.dumps(2048), # 限制单批次总token数 'LOAD_IN_4BIT': "true", # 启用4bit量化大幅降低显存占用 }
4bit量化可将Falcon-7B的显存占用从FP16的约13GB降至约4GB,直接缓解内存压力。
2. 针对模型类型不兼容问题:添加远程代码信任配置
Falcon模型依赖自定义代码实现,推至Hugging Face Hub后部署时,必须启用信任远程代码的配置:
config = { 'HF_MODEL_ID': "your-hf-username/your-falcon-model", 'SM_NUM_GPUS': json.dumps(number_of_gpu), 'MAX_INPUT_LENGTH': json.dumps(1024), 'MAX_TOTAL_TOKENS': json.dumps(2048), 'HF_TRUST_REMOTE_CODE': "true", # 必须添加,加载Falcon自定义代码 'LOAD_IN_4BIT': "true", }
3. 模型打包优化:确保压缩包结构正确
使用model.tar.gz部署时,需确保压缩包解压后直接包含模型核心文件(pytorch_model.bin、config.json等),而非嵌套子目录。打包命令示例:
cd /path/to/model-files tar -czf model.tar.gz *
上传至S3后,部署时需指定该压缩包的S3路径,而非S3前缀。
4. 实例资源优化:启用张量并行
如果使用多GPU实例,可启用张量并行进一步分摊内存压力:
config = { 'HF_MODEL_ID': "/opt/ml/model", 'SM_NUM_GPUS': json.dumps(2), 'MAX_INPUT_LENGTH': json.dumps(1024), 'MAX_TOTAL_TOKENS': json.dumps(2048), 'MAX_BATCH_PREFILL_TOKENS': json.dumps(1024), 'LOAD_IN_4BIT': "true", 'TP_SIZE': json.dumps(2), # 张量并行数与GPU数量一致 }
内容的提问来源于stack exchange,提问作者Jacob Brophy
相关产品推荐
相关产品推荐

