You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ECS Fargate触发CPU低利用率告警后未自动缩容问题

ECS自动缩容失效问题排查与修复

我有一个ECS集群,配置了两个基于CPU利用率的伸缩告警:CPU利用率达到60%时触发扩容策略,低于等于10%时触发缩容策略。扩容功能运行正常——使用Siege工具压测(CPU利用率超过60%)时,集群成功创建新任务,但停止所有流量后,CPU利用率降至10%以下,任务却未被自动删除。以下是我的自动扩缩容Terraform配置代码:

//create auto scaling target
resource "aws_appautoscaling_target" "rcc-ecs_target" {
  service_namespace  = "ecs"
  resource_id        = "service/${aws_ecs_cluster.rcc-cluster.name}/${aws_ecs_service.rcc-service.name}"
  scalable_dimension = "ecs:service:DesiredCount"
  min_capacity       = 1
  max_capacity       = 3
}

//create autoscaling policy "up"
resource "aws_appautoscaling_policy" "rcc-up" {
  name               = "scale_up"
  service_namespace  = aws_appautoscaling_target.rcc-ecs_target.service_namespace
  resource_id        = aws_appautoscaling_target.rcc-ecs_target.resource_id
  scalable_dimension = aws_appautoscaling_target.rcc-ecs_target.scalable_dimension

  step_scaling_policy_configuration {
    adjustment_type         = "ChangeInCapacity"
    cooldown                = 60
    metric_aggregation_type = "Maximum"

    step_adjustment {
      metric_interval_lower_bound = 0
      scaling_adjustment          = 1
    }
  }
  depends_on = [aws_appautoscaling_target.rcc-ecs_target]
}

//create autoscaling policy "down"
resource "aws_appautoscaling_policy" "rcc-down" {
  name               = "scale_down"
  service_namespace  = aws_appautoscaling_target.rcc-ecs_target.service_namespace
  resource_id        = aws_appautoscaling_target.rcc-ecs_target.resource_id
  scalable_dimension = aws_appautoscaling_target.rcc-ecs_target.scalable_dimension

  step_scaling_policy_configuration {
    adjustment_type         = "ChangeInCapacity"
    cooldown                = 60
    metric_aggregation_type = "Maximum"

    step_adjustment {
      metric_interval_lower_bound = 0
      scaling_adjustment          = -1
    }
  }
  depends_on = [aws_appautoscaling_target.rcc-ecs_target]
}

//create alarm that triggers the "up" policy
resource "aws_cloudwatch_metric_alarm" "rcc-cpu_high" {
  alarm_name          = "rcc-cpu_utilization_high"
  comparison_operator = "GreaterThanOrEqualToThreshold"
  evaluation_periods  = "1"
  metric_name         = "CPUUtilization"
  namespace           = "AWS/ECS"
  period              = "60"
  statistic           = "Average"
  threshold           = 60
  alarm_description   = "High CPU utilization"

  dimensions = {
    ClusterName = aws_ecs_cluster.rcc-cluster.name
    ServiceName = aws_ecs_service.rcc-service.name
  }
  alarm_actions = [aws_appautoscaling_policy.rcc-up.arn]
}

//create alarm that triggers the "down" policy
resource "aws_cloudwatch_metric_alarm" "rcc-cpu_low" {
  alarm_name          = "rcc-cpu_utilization_low"
  comparison_operator = "LessThanOrEqualToThreshold"
  evaluation_periods  = "1"
  metric_name         = "CPUUtilization"
  namespace           = "AWS/ECS"
  period              = "60"
  statistic           = "Average"
  threshold           = 10
  alarm_description   = "Low CPU utilization"

  dimensions = {
    ClusterName = aws_ecs_cluster.rcc-cluster.name
    ServiceName = aws_ecs_service.rcc-service.name
  }
  alarm_actions = [aws_appautoscaling_policy.rcc-down.arn]
}

问题根源

缩容策略的step_adjustment配置逻辑错误:

  1. 当前缩容策略设置metric_interval_lower_bound = 0,表示仅当指标值大于0时才会触发-1的容量调整,但缩容告警触发条件是CPU≤10,此时指标值可能落在0-10区间,无法匹配该调整规则。
  2. 缩放策略的metric_aggregation_type使用了Maximum,但CloudWatch告警的统计值是Average,统计方式不匹配可能导致策略触发异常。

修复方案

修改缩容策略的配置,调整step规则并对齐统计方式:

//create autoscaling policy "down"
resource "aws_appautoscaling_policy" "rcc-down" {
  name               = "scale_down"
  service_namespace  = aws_appautoscaling_target.rcc-ecs_target.service_namespace
  resource_id        = aws_appautoscaling_target.rcc-ecs_target.resource_id
  scalable_dimension = aws_appautoscaling_target.rcc-ecs_target.scalable_dimension

  step_scaling_policy_configuration {
    adjustment_type         = "ChangeInCapacity"
    cooldown                = 60
    metric_aggregation_type = "Average" # 和告警统计值保持一致

    step_adjustment {
      metric_interval_upper_bound = 0 # 匹配缩容告警的触发逻辑
      scaling_adjustment          = -1
    }
  }
  depends_on = [aws_appautoscaling_target.rcc-ecs_target]
}

额外检查点

  1. 确认ECS服务部署类型为REPLICA,DAEMON类型不支持自动扩缩容。
  2. 在CloudWatch控制台查看缩容告警的历史状态,确认其是否进入过ALARM状态。
  3. 在Application Auto Scaling控制台查看缩容策略的执行记录,排查是否存在错误信息。

内容的提问来源于stack exchange,提问作者marius

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 21:40:17