You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Amazon SageMaker Autopilot AutoML训练部署集成到Pipeline中?

将Amazon SageMaker Autopilot整合到SageMaker Pipeline工作流的实践方案

完全可以把Autopilot AutoML的训练和部署流程整合进SageMaker Pipeline,以下是实际项目中验证过的实现步骤:

核心思路

通过SageMaker Pipeline的AutoMLJobStep封装Autopilot训练任务,再添加步骤提取最优模型,最后串联部署流程,形成端到端的自动化工作流。

具体实现步骤

1. 定义Autopilot训练步骤

用AutoMLJobStep将Autopilot的训练配置封装成Pipeline的一个步骤,指定输入数据、任务类型、目标列等核心参数:

import sagemaker
from sagemaker.automl import AutoML
from sagemaker.workflow.pipeline import Pipeline
from sagemaker.workflow.steps import AutoMLJobStep
from sagemaker.workflow.parameters import ParameterString

# 定义可配置参数
input_data = ParameterString(name="InputDataUrl", default_value="s3://your-bucket/training-dataset/")
target_column = ParameterString(name="TargetColumn", default_value="prediction_label")

# 初始化AutoML实例
automl = AutoML(
    role=sagemaker.get_execution_role(),
    target_attribute_name=target_column,
    problem_type="BinaryClassification",  # 根据任务类型调整
    max_candidates=15,  # 控制候选模型数量,平衡训练时间与效果
    base_job_name="automl-pipeline-training"
)

# 创建Autopilot Pipeline步骤
automl_train_step = AutoMLJobStep(
    name="AutoML-Training",
    automl_job=automl,
    inputs=automl.inputs(
        s3_data_input_path=input_data,
        target_attribute_name=target_column
    )
)

2. 提取Autopilot最优模型

Autopilot训练完成后,通过Pipeline的步骤获取最优候选模型的模型数据和镜像地址,封装成SageMaker Model对象:

import boto3
from sagemaker.workflow.model_step import ModelStep
from sagemaker.model import Model

def fetch_best_automl_model(automl_job_name):
    sagemaker_client = boto3.client("sagemaker")
    job_details = sagemaker_client.describe_auto_ml_job(AutoMLJobName=automl_job_name)
    best_candidate = job_details["BestCandidate"]
    
    # 获取模型 artifacts 和推理镜像
    model_artifacts = best_candidate["CandidateProperties"]["CandidateArtifactLocations"]["ModelArtifacts"]
    inference_image = best_candidate["InferenceContainers"][0]["Image"]
    
    return Model(
        model_data=model_artifacts,
        role=sagemaker.get_execution_role(),
        image_uri=inference_image
    )

# 创建模型提取步骤
model_extract_step = ModelStep(
    name="Extract-Best-Model",
    model=fetch_best_automl_model(automl_train_step.properties.AutoMLJobName)
)

3. 添加模型部署步骤

将提取到的最优模型部署到SageMaker端点,作为Pipeline的最后一步:

from sagemaker.workflow.steps import ModelDeployStep

deploy_step = ModelDeployStep(
    name="Deploy-Model-To-Endpoint",
    model=model_extract_step.properties.ModelName,
    initial_instance_count=1,
    instance_type="ml.t2.medium",  # 根据模型大小和性能需求调整
    endpoint_name="automl-pipeline-endpoint"
)

4. 组装并启动Pipeline

将所有步骤串联成完整的工作流,提交执行:

# 构建Pipeline
pipeline = Pipeline(
    name="End-to-End-Autopilot-Pipeline",
    parameters=[input_data, target_column],
    steps=[automl_train_step, model_extract_step, deploy_step]
)

# 提交Pipeline
pipeline.upsert(role_arn=sagemaker.get_execution_role())
pipeline.start()

关键注意事项

  • 权限配置:确保Pipeline使用的IAM角色拥有Autopilot训练、模型部署、S3访问、CloudWatch日志等全流程权限
  • 流程扩展:可添加条件分支(比如根据模型精度阈值决定是否部署)、数据验证步骤、模型监控步骤,完善工作流
  • 参数调优:Autopilot的max_candidates、max_runtime_per_training_job_seconds等参数需根据数据集大小和任务需求调整,避免超时或资源浪费

内容的提问来源于stack exchange,提问作者Uwais Iqbal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 02:03:11