You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在SageMaker Pipeline UI中展示自定义任务的指标与数值?

在SageMaker Pipeline步骤输出标签页展示自定义容器指标

针对不同类型的Pipeline步骤,实现自定义容器指标展示的方式略有不同,具体操作如下:

一、TrainingStep(训练步骤)

自定义训练容器只需按规则输出指标,配合Estimator的配置即可自动捕获:

  • 在训练脚本中,将指标以固定格式打印到stdout或stderr,例如:
    print(f"epoch=3 train_loss=0.15 val_accuracy=0.94")
    
  • 定义Estimator时,通过metric_definitions参数设置正则表达式匹配指标:
    from sagemaker.estimator import Estimator
    
    estimator = Estimator(
        image_uri="你的自定义训练镜像地址",
        instance_type="ml.m5.xlarge",
        instance_count=1,
        metric_definitions=[
            {"Name": "train_loss", "Regex": "train_loss=([0-9\\.]+)"},
            {"Name": "val_accuracy", "Regex": "val_accuracy=([0-9\\.]+)"}
        ]
    )
    
    训练任务完成后,TrainingStep的输出标签页会自动展示捕获到的指标曲线和数值。

二、ProcessingStep(处理步骤)

需将指标写入指定路径的结构化文件,SageMaker会自动识别并展示:

  • 在自定义处理脚本中,将指标以JSON格式写入/opt/ml/output/metrics/目录下的文件(文件名可自定义,比如processing_metrics.json),格式示例:
    {
        "metrics": [
            {"name": "valid_data_ratio", "value": 99.2, "unit": "Percent"},
            {"name": "duplicate_record_count", "value": 47, "unit": "Count"}
        ]
    }
    
  • 确保自定义容器具备写入该目录的权限,ProcessingStep运行后,输出标签页会加载这些指标数据。

三、LambdaStep(Lambda步骤)

通过日志捕获或CloudWatch指标同步来展示:

  • 日志捕获方式:在Lambda函数中按固定格式打印日志,例如:
    import logging
    
    logger = logging.getLogger()
    logger.setLevel(logging.INFO)
    logger.info(f"lambda_processing: total_items=12000 success_rate=99.7")
    
    SageMaker会自动从Lambda的CloudWatch日志中提取匹配的指标,展示在输出标签页。
  • CloudMetric同步方式:在Lambda中调用CloudWatch的PutMetricData API,将指标发送到CloudWatch,之后可在Studio中关联查看这些指标。

内容的提问来源于stack exchange,提问作者PolarStorm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 02:50:51