如何在SageMaker Pipeline中集成拟合编码器模型部署Endpoint?
SageMaker Pipeline集成预训练预处理器模型的Step选择
针对你提到的场景——集成已拟合完成的预处理器模型到PipelineModel并部署推理端点,结论如下:
核心判断:是否需要重新训练预处理器
如果预处理器是离线拟合完成、无需在Pipeline中重新训练的模型:
不需要使用TrainingStep,直接通过以下方式集成:- 用
sagemaker.model.Model类加载已有的预处理器模型(指定模型存储路径、对应的推理镜像)。 - 将加载后的预处理器模型,与XGBoost训练步骤(如果XGBoost需要在Pipeline中训练则用
TrainingStep)输出的模型一起,传入PipelineModel组合成推理管道。 - 后续通过
CreateModelStep创建组合后的模型实体,再配合EndpointConfigStep和CreateEndpointStep完成端点部署。
- 用
如果预处理器需要在Pipeline运行时基于新数据重新拟合训练:
这种情况才需要使用TrainingStep来执行预处理器的训练流程,把训练输出的模型产物作为PipelineModel的输入组件之一。
适配你的场景的操作示例片段
# 加载已拟合好的预处理器模型 preprocessor_model = Model( model_data="s3://your-bucket/path/to/trained-preprocessor/model.tar.gz", image_uri=sklearn_inference_image_uri, role=sagemaker_role ) # 假设XGBoost模型是通过TrainingStep训练得到的 xgboost_model = TrainingStep( name="XGBoostTraining", estimator=xgb_estimator, inputs=TrainingInput(s3_data="s3://your-bucket/train-data") ) # 构建PipelineModel pipeline_model = PipelineModel( name="Preprocessor-XGB-Pipeline", models=[preprocessor_model, xgboost_model], role=sagemaker_role ) # 创建模型实体的Step create_model_step = CreateModelStep( name="CreatePipelineModel", model=pipeline_model ) # 后续添加端点配置和创建端点的Step
内容的提问来源于stack exchange,提问作者soulwreckedyouth
相关产品推荐
相关产品推荐

