You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化SageMaker异步端点0到1实例扩容的延迟问题

问题描述

使用AWS SageMaker异步端点,目标是无请求时缩容至0实例以节省成本,有请求时从0扩容到1实例处理单请求。已配置基于HasBacklogWithoutCapacity指标的自动扩缩容策略(包括TargetTracking和StepScaling),但存在扩容启动时机延迟以及0到1实例扩容的latency问题。

当前配置详情:

  • SageMaker异步端点
  • 基于HasBacklogWithoutCapacity指标的自动扩缩容策略

TargetTracking扩缩容策略代码

scaling_policy = application_scaling.put_scaling_policy(
    PolicyName=policy_name,
    ServiceNamespace="sagemaker",  # The namespace of the service that provides the resource.
    ResourceId=resource_id,  # Endpoint name
    ScalableDimension="sagemaker:variant:DesiredInstanceCount",  
    PolicyType="TargetTrackingScaling",
    TargetTrackingScalingPolicyConfiguration={
        'TargetValue': 1.0,
        'CustomizedMetricSpecification': {
            'MetricName': 'HasBacklogWithoutCapacity',
            'Namespace': 'AWS/SageMaker',
            'Dimensions': [
                {'Name': 'EndpointName', 'Value': endpoint_name}
            ],
            'Statistic': 'Average',
        },
        'ScaleInCooldown': 60,
        'ScaleOutCooldown': 10,
    }
)

StepScaling扩缩容策略代码

scaleout_policy_name = f"{endpoint_name}-ScalingoutPolicy"

scalingout_policy = application_scaling.put_scaling_policy(
    PolicyName=scaleout_policy_name,
    ServiceNamespace="sagemaker",
    ResourceId=resource_id,
    ScalableDimension="sagemaker:variant:DesiredInstanceCount",
    PolicyType="StepScaling",
    StepScalingPolicyConfiguration={
        'AdjustmentType': 'ChangeInCapacity',
        'StepAdjustments': [
            {
                'MetricIntervalLowerBound': 0,
                # 'MetricIntervalUpperBound': 10.0,
                'ScalingAdjustment': 1  
            },

        ],
        'Cooldown': 30, 
        'MetricAggregationType': 'Average'
    }
)

scaleout_alarm_name = f"{endpoint_name}-ScalingoutAlarm"
scale_out_policy_arn = scalingout_policy['PolicyARN']

scalingout_alarm = cloudwatch_client.put_metric_alarm(
    AlarmName=scaleout_alarm_name,
    MetricName='HasBacklogWithoutCapacity',
    Namespace='AWS/SageMaker',
    Statistic='Average',
    Period=60,  # Period in seconds
    EvaluationPeriods=1,
    Threshold=1.0, 
    ComparisonOperator='GreaterThanOrEqualToThreshold',
    Dimensions=[
        {
            'Name': 'EndpointName',
            'Value': endpoint_name
        },
        {
            'Name': 'VariantName',
            'Value': 'AllTraffic'
        }
    ],
    AlarmActions=[
        scale_out_policy_arn
    ],
    AlarmDescription='Alarm for scaling out SageMaker endpoint instances',
    Unit='Count'
)
优化方案与建议

1. 缩小CloudWatch指标统计周期

当前StepScaling配置的Period=60秒是启动延迟的核心因素之一,CloudWatch每60秒才聚合一次指标,导致告警触发滞后。将周期改为10秒,能让系统更快检测到待处理请求,触发扩容:

# 修改StepScaling告警的Period参数
scalingout_alarm = cloudwatch_client.put_metric_alarm(
    # 其他参数保持不变
    Period=10,  # 从60秒调整为10秒
    EvaluationPeriods=1,
    # 其他参数保持不变
)

对于TargetTracking策略,可通过Lambda自定义实时上报Backlog状态的CloudWatch指标,进一步缩短检测间隔。

2. 优化扩缩容冷却时间

  • 针对0到1的扩容场景,将ScaleOutCooldown(TargetTracking)或Cooldown(StepScaling)设置为5秒,避免冷却时间阻碍快速扩容(0到1场景下重复扩容风险极低);
  • 保持ScaleInCooldown在60秒以上,避免实例刚启动就被误缩容。

3. 替换扩容触发指标

HasBacklogWithoutCapacity是布尔型指标(0或1),平均统计可能延迟触发。改用PendingRequests指标直接监控待处理请求数,触发逻辑更精准:

TargetTracking策略调整:

'TargetTrackingScalingPolicyConfiguration': {
    'TargetValue': 0.0,  # 目标待处理请求数为0
    'CustomizedMetricSpecification': {
        'MetricName': 'PendingRequests',
        'Namespace': 'AWS/SageMaker',
        'Dimensions': [
            {'Name': 'EndpointName', 'Value': endpoint_name},
            {'Name': 'VariantName', 'Value': 'AllTraffic'}
        ],
        'Statistic': 'Maximum',  # 用最大值确保不遗漏请求
    },
    'ScaleOutCooldown': 5,
    'ScaleInCooldown': 120,
}

StepScaling策略调整:

修改告警的MetricName为PendingRequests,Threshold=1,ComparisonOperator='GreaterThanOrEqualToThreshold',只要有1个待处理请求就触发扩容。

4. 优化模型启动速度

如果实例启动后模型初始化耗时久,可从镜像层面优化:

  • 使用更小的基础镜像,减少镜像拉取时间;
  • 预加载模型权重到镜像中,避免启动时再下载;
  • 若预算允许,可搭配SageMaker弹性推理的预热基础实例,降低冷启动延迟。

5. 直接触发扩容(极端低延迟场景)

对于低频率但要求极低延迟的请求,可在发送请求前通过API直接设置实例数为1,处理完成后再缩容,绕过自动扩缩容的检测延迟:

# 发送请求前手动扩容
sagemaker_client.update_endpoint(
    EndpointName=endpoint_name,
    DesiredInstanceCount=1
)
# 等待实例就绪后发送请求
# 处理完成后手动缩容
sagemaker_client.update_endpoint(
    EndpointName=endpoint_name,
    DesiredInstanceCount=0
)

内容的提问来源于stack exchange,提问作者YUVAL MEHTA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 06:12:38