如何用Terraform配置AWS ECS EC2集群ASG基于服务副本数自动扩缩容
你的当前配置存在两个核心问题:
- 仅定义了ASG扩缩容策略,但没有绑定CloudWatch告警触发条件,策略永远不会被执行
- 试图基于服务副本数扩缩容的思路有偏差——副本数增加不代表现有ECS实例资源不足以承载,正确的触发指标应该是待调度任务数(Pending Tasks)或集群剩余资源容量,避免不必要的扩容。
一、核心思路调整
放弃直接基于服务副本数扩缩容,改用以下更合理的触发逻辑:
- 扩容触发:当ECS集群中存在待调度任务(Pending Tasks > 0),或集群CPU/内存利用率超过阈值(比如70%)时,自动扩容ASG
- 缩容触发:当ECS实例上无运行任务,且集群整体资源利用率低于阈值(比如30%)时,自动缩容ASG
二、完整Terraform配置实现
1. 定义ECS集群扩容触发的CloudWatch告警
# 扩容告警:当ECS集群待调度任务数持续5分钟大于0时触发 resource "aws_cloudwatch_metric_alarm" "ecs_pending_tasks_high" { alarm_name = "ecs-pending-tasks-high" comparison_operator = "GreaterThanThreshold" evaluation_periods = "2" metric_name = "PendingTasks" namespace = "AWS/ECS" period = "300" statistic = "Average" threshold = "0" alarm_description = "Trigger ASG scale-out when ECS has pending tasks" alarm_actions = [aws_autoscaling_policy.asg_scale_out.arn] dimensions = { ClusterName = aws_ecs_cluster.ecs_cluster.name # 替换为你的ECS集群名称 } }
2. 定义ASG扩容/缩容策略
# ASG扩容策略:每次增加1台实例 resource "aws_autoscaling_policy" "asg_scale_out" { name = "asg-scale-out" scaling_adjustment = 1 adjustment_type = "ChangeInCapacity" cooldown = 300 # 调整为合理冷却时间,避免频繁扩容 autoscaling_group_name = aws_autoscaling_group.ecs_asg.name policy_type = "SimpleScaling" } # ASG缩容策略:每次减少1台实例 resource "aws_autoscaling_policy" "asg_scale_in" { name = "asg-scale-in" scaling_adjustment = -1 adjustment_type = "ChangeInCapacity" cooldown = 300 autoscaling_group_name = aws_autoscaling_group.ecs_asg.name policy_type = "SimpleScaling" }
3. 精准缩容:基于ECS实例空闲状态(CloudWatch+Lambda)
要精准终止无任务的实例,推荐用CloudWatch事件触发Lambda函数,识别空闲实例后调用ASG终止接口:
# CloudWatch事件规则:每5分钟检查一次ECS实例状态 resource "aws_cloudwatch_event_rule" "ecs_idle_instance_check" { name = "ecs-idle-instance-check" description = "Check for idle ECS instances every 5 minutes" schedule_expression = "rate(5 minutes)" } # 事件目标:触发Lambda函数(需提前创建该Lambda) resource "aws_cloudwatch_event_target" "lambda_target" { rule = aws_cloudwatch_event_rule.ecs_idle_instance_check.name target_id = "lambda-idle-instance-terminator" arn = aws_lambda_function.ecs_idle_terminator.arn # 替换为你的Lambda ARN }
Lambda核心逻辑示例(Python):
import boto3 ecs = boto3.client('ecs') asg = boto3.client('autoscaling') def lambda_handler(event, context): cluster_name = "your-ecs-cluster-name" # 获取集群中所有ECS实例 instances = ecs.list_container_instances(cluster=cluster_name)['containerInstanceArns'] for instance_arn in instances: # 获取实例上的运行任务数 instance_details = ecs.describe_container_instances(cluster=cluster_name, containerInstances=[instance_arn])['containerInstances'][0] running_tasks = instance_details['runningTasksCount'] if running_tasks == 0: # 获取ECS实例对应的EC2实例ID并终止 ec2_instance_id = instance_details['ec2InstanceId'] asg.terminate_instance_in_auto_scaling_group( InstanceId=ec2_instance_id, ShouldDecrementDesiredCapacity=True )
4. 优化ASG基础配置
修改你的ASG资源,调整终止策略和健康检查逻辑:
resource "aws_autoscaling_group" "ecs_asg" { name = "ecs-asg" vpc_zone_identifier = [aws_subnet.public_1.id, aws_subnet.public_2.id, aws_subnet.public_3.id] launch_configuration = aws_launch_configuration.ecs_launch_config.name desired_capacity = 6 min_size = 1 # 不建议设为0,避免重启实例的调度延迟 max_size = 10 health_check_grace_period = 300 # 延长宽限期,给ECS Agent足够注册时间 health_check_type = "EC2" # 改用EC2健康检查,避免ELB逻辑干扰ECS实例状态 force_delete = true target_group_arns = [aws_lb_target_group.asg_tg.arn] termination_policies = ["NewestInstance"] # 优先终止最新实例,减少对运行服务的影响 }
三、额外建议
- 配置ECS实例保护:对运行关键任务的实例开启实例保护,避免被误终止
- 监控扩缩容行为:在CloudWatch中创建ASG容量变化告警,及时发现异常扩缩容
- 调整冷却时间:根据你的服务启动速度,合理设置ASG扩缩容的冷却时间,避免频繁波动
内容的提问来源于stack exchange,提问作者darian
相关产品推荐
相关产品推荐

