Sagemaker Pipelines部署PyTorch模型遇Endpoint超时及TrainingStep参数错误
SageMaker Pipelines部署PyTorch模型端点调用超时及TrainingStep参数错误问题
在SageMaker Studio中使用SageMaker Pipelines,通过PyTorch容器注册并部署自定义模型后,调用invoke_endpoint时端点出现超时错误:
ReadTimeoutError: Read timeout on endpoint URL: "https://runtime.sagemaker.eu-west-1.amazonaws.com/endpoints/nba-vw-base-endpoint-TEST/invocations"
检查端点日志未发现任何错误。
相关代码片段
模型训练与注册代码
##### PYTORCH CONTAINER # Step 1: Train Model # create model training instance model = PyTorch( entry_point="inference.py", framework_version='1.13', py_version='py39', source_dir="code", # sagemaker_session=pipeline_session, # I've tried this but doesn't work role=role, instance_type=training_instance, instance_count=1, base_job_name=f"{base_job_prefix}-{training_job_name}", output_path=s3_output_path, code_location=s3_training_output_path, # script_mode=True, hyperparameters={ "model_name": model_name, "model_type": model_type, "bucket": bucket, 'epsilon': 0.3 }, model_name=model_name + workflow_time ) # put it on the outside because fitting it inside TrainingStep isn't work model.fit() step_train = TrainingStep( name=training_step_name, # step_args=model.fit(), # I've tried this but it fails estimator=model, ) # Step 2: Register Model to Model Registry logger.info('Registering to model to Model Registry') step_register = RegisterModel( name=register_model_step_name, estimator=model, # model_data=step_train.properties.ModelArtifacts.S3ModelArtifacts, content_types=["application/json"], response_types=["application/json"], inference_instances=inference_instances, model_package_group_name=model_package_group_name, approval_status=model_approval_status, depends_on=[training_step_name] )
端点部署代码
# create an endpoint using model registry model config previosly created sm_client = boto3.client('sagemaker', region_name=AWS_REGION) create_endpoint_response = sm_client.create_endpoint( EndpointName=ENDPOINT_NAME, EndpointConfigName=endpoint_config_name )
已尝试的解决方案
- 使用Pipeline Session
- 在TrainingStep内部、外部调用
.fit(),或使用estimator参数 - 使用
RegisterModel()或model.register()
以上方案均未解决问题。参考官方示例时,若不使用pipeline_session调用model.fit(),会提示TrainingStep()需要estimator或step_args参数,说明.fit()未返回值。
更新:TrainingStep传入model.fit()的错误
尝试将.fit()放入TrainingStep中:
step_train = TrainingStep( name=training_step_name, step_args=model.fit(), )
训练任务日志显示已完成:
2024-02-14 13:16:00 Completed - Training job completed Training seconds: 112 Billable seconds: 112
但随即出现如下错误:
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) Cell In[22], line 1 ----> 1 step_train = TrainingStep( 2 name=training_step_name, 3 step_args=model.fit(), # need to fit the model to ensure it properly trains and creates inference logic 4 # estimator=model, # seems to be getting deprecated in future 5 ) File /opt/conda/lib/python3.10/site-packages/sagemaker/workflow/steps.py:417, in TrainingStep.__init__(self, name, step_args, estimator, display_name, description, inputs, cache_config, depends_on, retry_policies) 412 super(TrainingStep, self).__init__( 413 name, StepTypeEnum.TRAINING, display_name, description, depends_on, retry_policies 414 ) 416 if not (step_args is not None) ^ (estimator is not None): --> 417 raise ValueError("Either step_args or estimator need to be given.") 419 if step_args: 420 from sagemaker.workflow.utilities import validate_step_args_input ValueError: Either step_args or estimator need to be given.
这说明.fit()未返回值,导致传入None。不清楚为何官方示例未出现此问题,希望得到下一步排查方向。
内容的提问来源于stack exchange,提问作者Cris Pineda
相关产品推荐
相关产品推荐

