You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Terraform配置AWS ECS EC2集群ASG基于服务副本数自动扩缩容

你的当前配置存在两个核心问题:

  1. 仅定义了ASG扩缩容策略,但没有绑定CloudWatch告警触发条件,策略永远不会被执行
  2. 试图基于服务副本数扩缩容的思路有偏差——副本数增加不代表现有ECS实例资源不足以承载,正确的触发指标应该是待调度任务数(Pending Tasks)或集群剩余资源容量,避免不必要的扩容。

一、核心思路调整

放弃直接基于服务副本数扩缩容,改用以下更合理的触发逻辑:

  • 扩容触发:当ECS集群中存在待调度任务(Pending Tasks > 0),或集群CPU/内存利用率超过阈值(比如70%)时,自动扩容ASG
  • 缩容触发:当ECS实例上无运行任务,且集群整体资源利用率低于阈值(比如30%)时,自动缩容ASG

二、完整Terraform配置实现

1. 定义ECS集群扩容触发的CloudWatch告警

# 扩容告警:当ECS集群待调度任务数持续5分钟大于0时触发
resource "aws_cloudwatch_metric_alarm" "ecs_pending_tasks_high" {
  alarm_name          = "ecs-pending-tasks-high"
  comparison_operator = "GreaterThanThreshold"
  evaluation_periods  = "2"
  metric_name         = "PendingTasks"
  namespace           = "AWS/ECS"
  period              = "300"
  statistic           = "Average"
  threshold           = "0"
  alarm_description   = "Trigger ASG scale-out when ECS has pending tasks"
  alarm_actions       = [aws_autoscaling_policy.asg_scale_out.arn]

  dimensions = {
    ClusterName = aws_ecs_cluster.ecs_cluster.name # 替换为你的ECS集群名称
  }
}

2. 定义ASG扩容/缩容策略

# ASG扩容策略:每次增加1台实例
resource "aws_autoscaling_policy" "asg_scale_out" {
  name                   = "asg-scale-out"
  scaling_adjustment     = 1
  adjustment_type        = "ChangeInCapacity"
  cooldown               = 300 # 调整为合理冷却时间,避免频繁扩容
  autoscaling_group_name = aws_autoscaling_group.ecs_asg.name
  policy_type            = "SimpleScaling"
}

# ASG缩容策略:每次减少1台实例
resource "aws_autoscaling_policy" "asg_scale_in" {
  name                   = "asg-scale-in"
  scaling_adjustment     = -1
  adjustment_type        = "ChangeInCapacity"
  cooldown               = 300
  autoscaling_group_name = aws_autoscaling_group.ecs_asg.name
  policy_type            = "SimpleScaling"
}

3. 精准缩容:基于ECS实例空闲状态(CloudWatch+Lambda)

要精准终止无任务的实例,推荐用CloudWatch事件触发Lambda函数,识别空闲实例后调用ASG终止接口:

# CloudWatch事件规则:每5分钟检查一次ECS实例状态
resource "aws_cloudwatch_event_rule" "ecs_idle_instance_check" {
  name        = "ecs-idle-instance-check"
  description = "Check for idle ECS instances every 5 minutes"
  schedule_expression = "rate(5 minutes)"
}

# 事件目标:触发Lambda函数(需提前创建该Lambda)
resource "aws_cloudwatch_event_target" "lambda_target" {
  rule      = aws_cloudwatch_event_rule.ecs_idle_instance_check.name
  target_id = "lambda-idle-instance-terminator"
  arn       = aws_lambda_function.ecs_idle_terminator.arn # 替换为你的Lambda ARN
}

Lambda核心逻辑示例(Python):

import boto3

ecs = boto3.client('ecs')
asg = boto3.client('autoscaling')

def lambda_handler(event, context):
    cluster_name = "your-ecs-cluster-name"
    # 获取集群中所有ECS实例
    instances = ecs.list_container_instances(cluster=cluster_name)['containerInstanceArns']
    for instance_arn in instances:
        # 获取实例上的运行任务数
        instance_details = ecs.describe_container_instances(cluster=cluster_name, containerInstances=[instance_arn])['containerInstances'][0]
        running_tasks = instance_details['runningTasksCount']
        if running_tasks == 0:
            # 获取ECS实例对应的EC2实例ID并终止
            ec2_instance_id = instance_details['ec2InstanceId']
            asg.terminate_instance_in_auto_scaling_group(
                InstanceId=ec2_instance_id,
                ShouldDecrementDesiredCapacity=True
            )

4. 优化ASG基础配置

修改你的ASG资源,调整终止策略和健康检查逻辑:

resource "aws_autoscaling_group" "ecs_asg" {
  name                      = "ecs-asg"
  vpc_zone_identifier       = [aws_subnet.public_1.id, aws_subnet.public_2.id, aws_subnet.public_3.id]
  launch_configuration      = aws_launch_configuration.ecs_launch_config.name
  desired_capacity          = 6
  min_size                  = 1 # 不建议设为0,避免重启实例的调度延迟
  max_size                  = 10
  health_check_grace_period = 300 # 延长宽限期,给ECS Agent足够注册时间
  health_check_type         = "EC2" # 改用EC2健康检查,避免ELB逻辑干扰ECS实例状态
  force_delete              = true
  target_group_arns         = [aws_lb_target_group.asg_tg.arn]
  termination_policies      = ["NewestInstance"] # 优先终止最新实例,减少对运行服务的影响
}

三、额外建议

  • 配置ECS实例保护:对运行关键任务的实例开启实例保护,避免被误终止
  • 监控扩缩容行为:在CloudWatch中创建ASG容量变化告警,及时发现异常扩缩容
  • 调整冷却时间:根据你的服务启动速度,合理设置ASG扩缩容的冷却时间,避免频繁波动

内容的提问来源于stack exchange,提问作者darian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 10:36:59