You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Terraform创建的AWS ECS自动扩缩容CloudWatch告警偶发失败

解决AWS ECS CloudWatch扩缩容告警的"No step adjustment found"错误

问题原因

你遇到的这个错误,核心是你的step scaling策略没有覆盖所有可能的指标值与阈值的偏差范围。结合你的配置来看:

  • 扩容告警设置了evaluation_periods=2(评估2个5分钟周期,共10分钟)和datapoints_to_alarm=1(只要1个数据点超标就触发)
  • 当触发告警时,Auto Scaling会同时评估这两个周期的指标值——其中一个是超过阈值500的516.95,另一个是低于阈值的437.08
  • 对于低于阈值的那个数据点,它与阈值的偏差是负数(437.08 - 500 = -62.91),但你的扩容策略里只定义了metric_interval_lower_bound=0(对应偏差>0的情况),完全没有处理偏差<0的场景,导致Auto Scaling找不到对应的调整步骤,从而抛出错误。

解决方案

你需要在step scaling策略中补充覆盖所有偏差范围的调整规则:

  • 对于扩容策略:添加一个处理偏差小于0的步骤,设置scaling_adjustment=0(也就是当指标值低于阈值时,不执行扩容操作)
  • 同理,缩容策略也需要补充处理偏差大于0的步骤,避免类似的错误发生

修改后的Terraform代码

扩容策略(task_count_up)

resource "aws_appautoscaling_policy" "task_count_up" {
  name               = "appScalingPolicy_${aws_ecs_service.sqs_to_kinesis.name}_ScaleUp"
  service_namespace  = "ecs"
  resource_id        = "service/${aws_ecs_cluster.shared-elb-access-logs-processor.name}/${aws_ecs_service.sqs_to_kinesis.name}"
  scalable_dimension = "ecs:service:DesiredCount"

  step_scaling_policy_configuration {
    adjustment_type          = "ChangeInCapacity"
    cooldown                 = "${var.scale_up_cooldown_seconds}"
    metric_aggregation_type  = "Maximum"

    # 处理指标值超过阈值的情况(偏差>0),扩容1个任务
    step_adjustment {
      metric_interval_lower_bound = 0
      scaling_adjustment          = 1
    }

    # 处理指标值低于阈值的情况(偏差<0),不进行任何调整
    step_adjustment {
      metric_interval_upper_bound = 0
      scaling_adjustment          = 0
    }
  }

  depends_on = [aws_appautoscaling_target.main]
}

缩容策略(task_count_down)

resource "aws_appautoscaling_policy" "task_count_down" {
  name               = "appScalingPolicy_${aws_ecs_service.sqs_to_kinesis.name}_ScaleDown"
  service_namespace  = "ecs"
  resource_id        = "service/${aws_ecs_cluster.shared-elb-access-logs-processor.name}/${aws_ecs_service.sqs_to_kinesis.name}"
  scalable_dimension = "ecs:service:DesiredCount"

  step_scaling_policy_configuration {
    adjustment_type          = "ChangeInCapacity"
    cooldown                 = "${var.scale_down_cooldown_seconds}"
    metric_aggregation_type  = "Minimum"

    # 处理指标值低于阈值的情况(偏差<0),缩容1个任务
    step_adjustment {
      metric_interval_upper_bound = 0
      scaling_adjustment          = -1
    }

    # 处理指标值高于阈值的情况(偏差>0),不进行任何调整
    step_adjustment {
      metric_interval_lower_bound = 0
      scaling_adjustment          = 0
    }
  }

  depends_on = [aws_appautoscaling_target.main]
}

额外说明

  • metric_interval_lower_bound和metric_interval_upper_bound是相对于阈值的偏差值,比如阈值是500,那么metric_interval_lower_bound=0就代表指标值 > 500,metric_interval_upper_bound=0代表指标值 < 500
  • 添加scaling_adjustment=0的步骤,就是告诉Auto Scaling当指标值处于这个范围时,不执行任何扩缩容操作,这样就能覆盖所有可能的指标值场景,避免出现找不到调整步骤的错误

内容的提问来源于stack exchange,提问作者vidya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:43:11