AWS CDK应用重部署时无法启动新EC2实例的问题求助
问题背景
我有一个通过GitHub Action频繁重部署的AWS CDK应用,基于EC2自动扩缩容组(ASG)的ECS服务运行。服务任务配置要求8192 CPU单位、30000MB内存,适配t2.2xlarge实例。重部署时预期流程是:AWS检查可用实例→发现旧实例资源不足→启动新实例→在新实例部署服务→终止旧实例。但实际报错:service was unable to place a task because no container instance met all of its requirements. The closest matching container-instance has insufficient CPU units available
需求是强制服务部署到新实例,而非旧实例。
原CDK配置
cluster = ecs.Cluster( self, "cluster", vpc=vpc, ) task_definition = ecs.Ec2TaskDefinition( self, "taskDefinition", execution_role=role, ) task_definition.add_container( "container", image=docker_image, port_mappings=[port_mapping], cpu=8192, memory_limit_mib=30_000, ) autoscaling_group = autoscaling.AutoScalingGroup( self, "autoscalingGroup", instance_type=ec2.InstanceType("t2.2xlarge"), min_capacity=1, max_capacity=2, update_policy=autoscaling.UpdatePolicy.rolling_update( min_instances_in_service=1, ), ) capacity_provider = ecs.AsgCapacityProvider( self, "asgCapacityProvider", auto_scaling_group=autoscaling_group, ) cluster.add_asg_capacity_provider(capacity_provider) service = ecs.Ec2Service( self, "service", cluster=cluster, task_definition=task_definition, )
原GitHub Action部署命令
aws ecs \ update-service \ --cluster "${VAR_CLUSTER}" \ --service "${VAR_SERVICE}" \ --force-new-deployment
核心问题分析
旧实例已运行着占用全部资源的任务,ECS默认会优先尝试在现有实例部署新任务,导致资源不足报错。同时,ASG容量提供商的托管扩缩容默认未启用,ECS无法自动触发扩容;即便启用,任务放置失败触发扩容的延迟也会导致部署直接失败。
解决措施
1. 配置容量提供商的托管扩缩容
开启ASG容量提供商的managed_scaling,让ECS根据任务需求自动调整实例数量;关闭managed_termination_protection,允许旧实例在新任务就绪后被终止。
2. 调整ECS服务的部署策略
设置minimum_healthy_percent=100确保旧任务持续运行直到新任务就绪,maximum_percent=200允许ECS启动新任务触发扩容;添加放置策略优先选择最新启动的实例,强制任务部署到新实例。
3. 优化部署流程(可选)
在部署前主动触发ASG扩容,确保有足够实例可用,待实例加入ECS集群后再执行部署命令。
修改后的配置
修改后的CDK代码
cluster = ecs.Cluster( self, "cluster", vpc=vpc, ) task_definition = ecs.Ec2TaskDefinition( self, "taskDefinition", execution_role=role, ) task_definition.add_container( "container", image=docker_image, port_mappings=[port_mapping], cpu=8192, memory_limit_mib=30_000, ) autoscaling_group = autoscaling.AutoScalingGroup( self, "autoscalingGroup", instance_type=ec2.InstanceType("t2.2xlarge"), min_capacity=1, max_capacity=2, update_policy=autoscaling.UpdatePolicy.rolling_update( min_instances_in_service=1, ), ) capacity_provider = ecs.AsgCapacityProvider( self, "asgCapacityProvider", auto_scaling_group=autoscaling_group, # 开启托管扩缩容,目标容量100%,确保ECS按需扩容 managed_scaling=ecs.ManagedScaling( target_capacity_percent=100, minimum_scaling_step_size=1, maximum_scaling_step_size=1, ), # 关闭托管终止保护,允许旧实例被终止 managed_termination_protection=ecs.ManagedTerminationProtection.DISABLED, ) cluster.add_asg_capacity_provider(capacity_provider) service = ecs.Ec2Service( self, "service", cluster=cluster, task_definition=task_definition, # 配置ECS部署控制器,确保先启动新任务再终止旧任务 deployment_controller=ecs.DeploymentController(type=ecs.DeploymentControllerType.ECS), minimum_healthy_percent=100, maximum_percent=200, # 放置策略:优先选择最新启动的实例 placement_strategies=[ ecs.PlacementStrategy.prioritize("attribute:ecs.instance-launch-time", ascending=False) ], )
修改后的GitHub Action部署命令(可选)
# 主动将ASG扩容到最大容量,跳过冷却期 aws autoscaling set-desired-capacity \ --auto-scaling-group-name "${VAR_ASG_NAME}" \ --desired-capacity 2 \ --no-honor-cooldown # 轮询等待实例加入ECS集群并处于ACTIVE状态 until aws ecs list-container-instances --cluster "${VAR_CLUSTER}" --query 'containerInstanceArns' | grep -q "arn"; do sleep 10 done # 执行强制部署 aws ecs \ update-service \ --cluster "${VAR_CLUSTER}" \ --service "${VAR_SERVICE}" \ --force-new-deployment # 部署完成后可选缩容回最小容量 # aws autoscaling set-desired-capacity \ # --auto-scaling-group-name "${VAR_ASG_NAME}" \ # --desired-capacity 1
内容的提问来源于stack exchange,提问作者Faust

