如何在Amazon SageMaker Pipeline各步骤中获取Pipeline执行ID
在Amazon SageMaker Pipeline步骤中获取执行ID的解决方案
核心问题
SageMaker Pipeline的执行变量(如ExecutionVariables.PIPELINE_EXECUTION_ID)是延迟解析的占位符,无法直接转换为字符串使用,需通过特定方式在步骤中动态引用,以此规整存储各步骤的输出结果。
解决方案
1. 在内置步骤中直接引用执行变量
对于ClarifyCheckStep这类内置步骤,可直接将执行变量嵌入到输出路径或参数配置中,Pipeline执行时会自动替换为实际的执行ID字符串。示例代码如下:
from sagemaker.workflow.execution_variables import ExecutionVariables from sagemaker.clarify import DataConfig, BiasConfig, ClarifyCheckConfig from sagemaker.workflow.steps import ClarifyCheckStep import sagemaker session = sagemaker.Session() # 配置Clarify检查所需的参数 data_config = DataConfig( s3_data_input_path="s3://your-input-data-path/", s3_output_path=None, # 后续通过步骤参数指定带执行ID的路径 label="your-label-column" ) bias_config = BiasConfig( label_values_or_threshold=["positive"], facet_name="your-facet-column" ) check_config = ClarifyCheckConfig( skip_checks=None, check_thresholds={"BiasOverall": 0.1} ) # 构建包含执行ID的输出路径 execution_id = ExecutionVariables.PIPELINE_EXECUTION_ID output_path = f"s3://your-bucket/clarify-results/{execution_id}/" # 创建ClarifyCheckStep clarify_check_step = ClarifyCheckStep( name="Clarify-Bias-Check-Step", data_config=data_config, bias_config=bias_config, check_config=check_config, output_path=output_path, sagemaker_session=session )
2. 在自定义脚本中通过环境变量获取
如果使用ProcessingStep、TrainingStep等支持自定义脚本的步骤,可直接从环境变量中读取执行ID:
- Bash脚本示例:
#!/bin/bash PIPELINE_EXECUTION_ID=$SM_PIPELINE_EXECUTION_ID # 将结果存储到包含执行ID的路径 mkdir -p /opt/ml/output/results/$PIPELINE_EXECUTION_ID cp your-result-file /opt/ml/output/results/$PIPELINE_EXECUTION_ID/
- Python脚本示例:
import os import pandas as pd execution_id = os.environ["SM_PIPELINE_EXECUTION_ID"] # 读取数据并处理 data = pd.read_csv("/opt/ml/input/data/train/train.csv") # 保存结果到包含执行ID的路径 output_path = f"/opt/ml/output/results/{execution_id}/" os.makedirs(output_path, exist_ok=True) data.to_csv(f"{output_path}/processed_data.csv", index=False)
官方文档执行变量说明(翻译)
执行变量是SageMaker Pipeline提供的动态占位符,用于在Pipeline执行阶段注入运行时信息,常用变量包括:
ExecutionVariables.PIPELINE_NAME: 当前运行的Pipeline名称ExecutionVariables.PIPELINE_EXECUTION_ID: 当前Pipeline的唯一执行IDExecutionVariables.PIPELINE_EXECUTION_ARN: 当前Pipeline执行的ARN(亚马逊资源名称)ExecutionVariables.PIPELINE_PARAMETER_PREFIX: Pipeline参数对应的环境变量前缀
这些变量属于延迟解析的占位符,不能直接作为字符串操作,必须在支持动态变量的Pipeline组件中使用(如步骤的输出路径、参数配置、环境变量定义等),Pipeline执行时会自动将其替换为对应的实际值。
内容的提问来源于stack exchange,提问作者soulwreckedyouth
相关产品推荐
相关产品推荐

