Terraform创建的AWS ECS自动扩缩容CloudWatch告警偶发失败
解决AWS ECS CloudWatch扩缩容告警的"No step adjustment found"错误
问题原因
你遇到的这个错误,核心是你的step scaling策略没有覆盖所有可能的指标值与阈值的偏差范围。结合你的配置来看:
- 扩容告警设置了
evaluation_periods=2(评估2个5分钟周期,共10分钟)和datapoints_to_alarm=1(只要1个数据点超标就触发) - 当触发告警时,Auto Scaling会同时评估这两个周期的指标值——其中一个是超过阈值500的516.95,另一个是低于阈值的437.08
- 对于低于阈值的那个数据点,它与阈值的偏差是负数(437.08 - 500 = -62.91),但你的扩容策略里只定义了
metric_interval_lower_bound=0(对应偏差>0的情况),完全没有处理偏差<0的场景,导致Auto Scaling找不到对应的调整步骤,从而抛出错误。
解决方案
你需要在step scaling策略中补充覆盖所有偏差范围的调整规则:
- 对于扩容策略:添加一个处理偏差小于0的步骤,设置
scaling_adjustment=0(也就是当指标值低于阈值时,不执行扩容操作) - 同理,缩容策略也需要补充处理偏差大于0的步骤,避免类似的错误发生
修改后的Terraform代码
扩容策略(task_count_up)
resource "aws_appautoscaling_policy" "task_count_up" { name = "appScalingPolicy_${aws_ecs_service.sqs_to_kinesis.name}_ScaleUp" service_namespace = "ecs" resource_id = "service/${aws_ecs_cluster.shared-elb-access-logs-processor.name}/${aws_ecs_service.sqs_to_kinesis.name}" scalable_dimension = "ecs:service:DesiredCount" step_scaling_policy_configuration { adjustment_type = "ChangeInCapacity" cooldown = "${var.scale_up_cooldown_seconds}" metric_aggregation_type = "Maximum" # 处理指标值超过阈值的情况(偏差>0),扩容1个任务 step_adjustment { metric_interval_lower_bound = 0 scaling_adjustment = 1 } # 处理指标值低于阈值的情况(偏差<0),不进行任何调整 step_adjustment { metric_interval_upper_bound = 0 scaling_adjustment = 0 } } depends_on = [aws_appautoscaling_target.main] }
缩容策略(task_count_down)
resource "aws_appautoscaling_policy" "task_count_down" { name = "appScalingPolicy_${aws_ecs_service.sqs_to_kinesis.name}_ScaleDown" service_namespace = "ecs" resource_id = "service/${aws_ecs_cluster.shared-elb-access-logs-processor.name}/${aws_ecs_service.sqs_to_kinesis.name}" scalable_dimension = "ecs:service:DesiredCount" step_scaling_policy_configuration { adjustment_type = "ChangeInCapacity" cooldown = "${var.scale_down_cooldown_seconds}" metric_aggregation_type = "Minimum" # 处理指标值低于阈值的情况(偏差<0),缩容1个任务 step_adjustment { metric_interval_upper_bound = 0 scaling_adjustment = -1 } # 处理指标值高于阈值的情况(偏差>0),不进行任何调整 step_adjustment { metric_interval_lower_bound = 0 scaling_adjustment = 0 } } depends_on = [aws_appautoscaling_target.main] }
额外说明
metric_interval_lower_bound和metric_interval_upper_bound是相对于阈值的偏差值,比如阈值是500,那么metric_interval_lower_bound=0就代表指标值 > 500,metric_interval_upper_bound=0代表指标值 < 500- 添加
scaling_adjustment=0的步骤,就是告诉Auto Scaling当指标值处于这个范围时,不执行任何扩缩容操作,这样就能覆盖所有可能的指标值场景,避免出现找不到调整步骤的错误
内容的提问来源于stack exchange,提问作者vidya
相关产品推荐
相关产品推荐

