SageMaker Pipeline时间戳共享:控制GUI名称与S3存储路径
解决方案
1. 生成统一时间戳并跨文件传递
在run_pipeline.py中生成格式统一的UTC时间戳,用它设置Pipeline执行显示名,同时传递给pipeline.py的Pipeline构建函数,确保两端使用同一值:
# run_pipeline.py import time from pipeline import create_sagemaker_pipeline # 生成排序友好、易匹配的时间戳 time_stamp = time.strftime("%Y-%m-%d--%H-%M-%S", time.gmtime()) # 传递时间戳给Pipeline定义逻辑 pipeline = create_sagemaker_pipeline(execution_timestamp=time_stamp) # 启动Pipeline时用该时间戳作为显示名 execution = pipeline.start(execution_display_name=f"pipeline-exec-{time_stamp}")
2. 在Pipeline定义中复用时间戳配置S3路径
修改pipeline.py的构建函数,接收时间戳参数,并用它构建唯一的S3工件存储路径,确保所有Pipeline产出物都归类到对应时间戳的目录下:
# pipeline.py import boto3 from sagemaker.workflow.pipeline import Pipeline from sagemaker.workflow.steps import ProcessingStep from sagemaker.processing import ScriptProcessor, ProcessingInput, ProcessingOutput def create_sagemaker_pipeline(execution_timestamp): sagemaker_client = boto3.client("sagemaker") bucket = "your-s3-bucket-name" # 基于时间戳构建唯一S3根路径 s3_artifact_root = f"s3://{bucket}/sagemaker-pipeline-artifacts/{execution_timestamp}" # 示例:配置数据处理步骤的输出路径 processor = ScriptProcessor( command=["python3"], image_uri="your-processing-image-uri", role="your-sagemaker-role-arn", instance_count=1, instance_type="ml.t3.medium" ) processing_step = ProcessingStep( name="DataPreprocessing", processor=processor, inputs=[ProcessingInput(source="s3://your-input-data-path", destination="/opt/ml/processing/input")], outputs=[ProcessingOutput(source="/opt/ml/processing/output", destination=f"{s3_artifact_root}/preprocessed-data")], code="preprocess.py" ) # 构建Pipeline,将时间戳加入描述便于快速识别 pipeline = Pipeline( name="YourMLPipeline", steps=[processing_step], sagemaker_client=sagemaker_client, description=f"Execution timestamp: {execution_timestamp}" ) return pipeline
3. 灵活方案:用Pipeline参数动态传递时间戳
如果需要在启动执行阶段才确定时间戳(而非Pipeline构建阶段),可以通过PipelineParameter实现动态传递:
在pipeline.py中定义参数:
# pipeline.py from sagemaker.workflow.parameters import StringParameter def create_sagemaker_pipeline(): # 定义时间戳参数,启动时覆盖默认值 exec_timestamp = StringParameter(name="ExecutionTimestamp", default_value="temp-timestamp") bucket = "your-s3-bucket-name" # 用参数占位符拼接S3路径 s3_root = f"s3://{bucket}/sagemaker-pipeline-artifacts/{{{exec_timestamp.name}}}" # 后续步骤中使用该参数配置路径(示例同前) # ... pipeline = Pipeline( name="YourMLPipeline", parameters=[exec_timestamp], steps=[processing_step], # ... ) return pipeline
在run_pipeline.py中传递参数值:
# run_pipeline.py import time from pipeline import create_sagemaker_pipeline time_stamp = time.strftime("%Y-%m-%d--%H-%M-%S", time.gmtime()) pipeline = create_sagemaker_pipeline() # 启动时传递时间戳参数,同时设置显示名 execution = pipeline.start( execution_display_name=f"pipeline-exec-{time_stamp}", parameters={"ExecutionTimestamp": time_stamp} )
效果验证
- SageMaker控制台中,Pipeline执行记录的名称会包含时间戳,便于快速定位
- 所有Pipeline工件(预处理数据、模型、日志等)都会存储到对应时间戳的S3目录下,可直接匹配执行记录与工件文件
内容的提问来源于stack exchange,提问作者Francisco C
相关产品推荐
相关产品推荐

