AWS实例调度器搭配Auto Scaling Groups:自动停机防实例终止方案问询
解决AWS Instance Scheduler与Auto Scaling Group的冲突问题
我之前在开发环境里也碰到过完全一样的痛点——用Instance Scheduler自动停实例,结果ASG一检测到实例不健康就直接给终止了,手动暂停ASG完全违背了自动化调度的初衷。下面是几个我实际验证过的自动化解决方案,全程无需手动干预:
方案1:直接让Instance Scheduler管理Auto Scaling Group(最推荐)
如果你的核心需求是非工作时间停止所有实例、工作时间恢复,其实没必要单独调度EC2实例,直接把调度目标改成ASG就行:
- 删掉原来针对单个EC2实例的调度规则,新建针对目标ASG的调度配置。
- Instance Scheduler会在非工作时间自动将ASG的
DesiredCapacity和MinSize设为0,这样ASG不会尝试维持实例数量,自然不会触发不健康实例替换。 - 工作时间到来时,调度器会把这两个值恢复到你预设的数值,ASG会自动启动对应数量的实例(如果之前的实例只是停机没被终止,也会优先启动它们)。
这个方案最省心,完全利用Instance Scheduler的原生能力,不需要额外写代码。
方案2:用Lambda+EventBridge联动,动态暂停/恢复ASG的不健康替换进程
如果你的需求是保留停机的实例(而不是让ASG重新创建),可以通过自动化脚本在实例启停时调整ASG的进程:
步骤1:编写两个Lambda函数
暂停ASG的ReplaceUnhealthy进程(实例停机时触发)
import boto3 asg_client = boto3.client('autoscaling') def lambda_handler(event, context): # 从EventBridge事件中获取停机的实例ID instance_id = event['detail']['instance-id'] # 查询实例所属的ASG asg_response = asg_client.describe_auto_scaling_instances(InstanceIds=[instance_id]) if not asg_response['AutoScalingInstances']: return {"status": "success", "message": "Instance not part of any ASG"} asg_name = asg_response['AutoScalingInstances'][0]['AutoScalingGroupName'] # 暂停ReplaceUnhealthy进程 asg_client.suspend_processes( AutoScalingGroupName=asg_name, ScalingProcesses=['ReplaceUnhealthy'] ) return {"status": "success", "message": f"Suspended ReplaceUnhealthy for ASG {asg_name}"}
恢复ASG的ReplaceUnhealthy进程(实例启动时触发)
import boto3 asg_client = boto3.client('autoscaling') def lambda_handler(event, context): instance_id = event['detail']['instance-id'] asg_response = asg_client.describe_auto_scaling_instances(InstanceIds=[instance_id]) if not asg_response['AutoScalingInstances']: return {"status": "success", "message": "Instance not part of any ASG"} asg_name = asg_response['AutoScalingInstances'][0]['AutoScalingGroupName'] # 恢复ReplaceUnhealthy进程 asg_client.resume_processes( AutoScalingGroupName=asg_name, ScalingProcesses=['ReplaceUnhealthy'] ) return {"status": "success", "message": f"Resumed ReplaceUnhealthy for ASG {asg_name}"}
步骤2:配置EventBridge触发规则
- 停机触发规则:创建规则匹配
EC2 Instance State-change Notification,状态选择stopping,目标设置为上面的暂停Lambda。 - 启动触发规则:创建规则匹配
EC2 Instance State-change Notification,状态选择starting,目标设置为恢复Lambda。
步骤3:给Lambda配置权限
确保Lambda的执行角色拥有以下权限:
autoscaling:DescribeAutoScalingInstancesautoscaling:SuspendProcessesautoscaling:ResumeProcesses
方案3:自定义ASG健康检查逻辑
如果不想修改ASG进程,也可以调整健康检查的触发条件:
- 把ASG的EC2健康检查
HealthCheckGracePeriod设置为非工作时间的总时长(比如12小时),这样在停机期间ASG不会立即判定实例不健康。 - 或者使用CloudWatch告警作为ASG的自定义健康检查,只在工作时间启用告警规则;非工作时间自动禁用告警,ASG就不会收到不健康信号。
这个方案灵活性稍差,适合非工作时间固定的场景。
内容的提问来源于stack exchange,提问作者hhh0505
相关产品推荐
相关产品推荐

