在Amazon SageMaker部署Bloom模型遇400错误的解决方法咨询
问题
我希望在Amazon SageMaker上部署Bloom模型,以获得可使用的Bloom推理API,在SageMaker Jupyter Notebook中运行以下代码:
from sagemaker.huggingface import HuggingFaceModel import sagemaker role = sagemaker.get_execution_role() # Hub Model configuration hub = { 'HF_MODEL_ID':'bigscience/bloom', 'HF_TASK':'text-generation' } # create Hugging Face Model Class huggingface_model = HuggingFaceModel( transformers_version='4.17.0', pytorch_version='1.10.2', py_version='py38', env=hub, role=role, ) # deploy model to SageMaker Inference predictor = huggingface_model.deploy( initial_instance_count=1, # number of instances instance_type='ml.m5.xlarge' # ec2 instance type ) predictor.predict({ 'inputs': "Can you please let us know more details about your " })
执行后出现如下错误:
--------------------------------------------------------------------------- ModelError Traceback (most recent call last) /tmp/ipykernel_15151/842216467.py in <cell line: 1>() ----> 1 predictor.predict({ 2 'inputs': "Can you please let us know more details about your " 3 }) ~/anaconda3/envs/python3/lib/python3.8/site-packages/sagemaker/predictor.py in predict(self, data, initial_args, target_model, target_variant, inference_id) 159 data, initial_args, target_model, target_variant, inference_id 160 ) --> 161 response = self.sagemaker_session.sagemaker_runtime_client.invoke_endpoint(**request_args) 162 return self._handle_response(response) 163 ~/anaconda3/envs/python3/lib/python3.8/site-packages/botocore/client.py in _api_call(self, *args, **kwargs) 393 "%s() only accepts keyword arguments." % py_operation_name) 394 # The "self" in this scope is referring to the BaseClient. --> 395 return self._make_api_call(operation_name, kwargs) 396 397 _api_call.__name__ = str(py_operation_name) ~/anaconda3/envs/python3/lib/python3.8/site-packages/botocore/client.py in _make_api_call(self, operation_name, api_params) 723 error_code = parsed_response.get("Error", {}).get("Code") 724 error_class = self.exceptions.from_code(error_code) --> 725 raise error_class(parsed_response, operation_name) 726 else: 727 return parsed_response ModelError: An error occurred (ModelError) when calling the InvokeEndpoint operation: Received client error (400) from primary with message "{ "code": 400, "type": "InternalServerException", "message": "'bloom'" } ". See https://us-east-1.console.aws.amazon.com/cloudwatch/home?region=us-east-1#logEventViewer:group=/aws/sagemaker/Endpoints/huggingface-pytorch-inference-2022-07-29-23-06-38-076 in account 162923941922 for more information.
CloudWatch日志仅显示:
2022-07-29T23:09:09,135 [INFO ] W-bigscience__bloom-4-stdout com.amazonaws.ml.mms.wlm.WorkerLifeCycle - raise PredictionException(str(e), 400)
解决方案
1. 升级Transformers与PyTorch版本
你当前使用的Transformers 4.17.0版本对Bloom模型的支持不完善,Bloom模型在Transformers 4.20.0及以上版本才得到完整适配。建议更新为以下兼容版本:
huggingface_model = HuggingFaceModel( transformers_version='4.26.0', pytorch_version='1.13.1', py_version='py39', env=hub, role=role, )
2. 更换适配的实例类型
Bloom模型参数量极大(基础版bigscience/bloom为1760亿参数),ml.m5.xlarge的内存与算力完全无法支撑模型加载与运行。根据模型大小选择对应实例:
- 若部署轻量版
bigscience/bloom-560m:可使用ml.g4dn.xlarge或ml.p3.2xlarge - 若部署完整
bigscience/bloom:需使用ml.p4d.24xlarge这类超大显存实例,或启用模型并行方案
3. 补充推理参数(可选)
在预测调用中添加文本生成相关参数,避免因参数缺失引发错误:
predictor.predict({ 'inputs': "Can you please let us know more details about your ", 'parameters': { 'max_new_tokens': 50, 'temperature': 0.7 } })
4. 启用模型并行(针对大参数量模型)
对于超大规模Bloom模型,可通过设置环境变量开启模型并行,将模型拆分到多个GPU上运行:
hub = { 'HF_MODEL_ID':'bigscience/bloom', 'HF_TASK':'text-generation', 'HF_MODEL_PARALLEL': 'true' # 开启模型并行 }
内容的提问来源于stack exchange,提问作者The Doctor
相关产品推荐
相关产品推荐

