EC2实例结合ALB与ASG异常终止重建问题排查求助
我搭建了两台EC2服务器,关联到ASG(自动伸缩组)并挂载到ALB(应用负载均衡),分别位于eu-west-2a和eu-west-2b可用区。这些新服务器持续被终止并重建,无法定位原因。
相关配置
ALB目标组配置
resource "aws_lb_target_group" "targetgroup" { name = "targetgroup" target_type = "instance" port = 80 protocol = "HTTP" vpc_id = aws_vpc.main_vpc.id health_check { path = "/" port = 80 protocol = "HTTP" healthy_threshold = 10 unhealthy_threshold = 10 matcher = "200-499" } }
安全组已开放80端口。
ASG活动错误信息
- 其中一条错误:
Terminating EC2 instance: i-04dbb4d0f8c06b355 - Waiting For ELB Connection Draining. At 2024-03-11T15:02:16Z an instance was taken out of service in response to an EC2 health check indicating it has been terminated or stopped.
- 更常见的错误:
Launching a new EC2 instance: i-0d4840406c9612b45. Status Reason: Instance became unhealthy while waiting for instance to be in InService state. Termination Reason: Client.InstanceInitiatedShutdown: Instance initiated shutdown
ASG配置
resource "aws_autoscaling_group" "my_asg" { name = "my_asg" max_size = 1 min_size = 1 health_check_type = "ELB" # optional desired_capacity = 1 target_group_arns = [aws_lb_target_group.targetgroup.arn] health_check_grace_period = 300 vpc_zone_identifier = [aws_subnet.public-subnet-1.id, aws_subnet.public-subnet-2.id] }
伸缩策略配置
# ASG扩容策略 resource "aws_autoscaling_policy" "scale_up" { name = "scale_up" policy_type = "SimpleScaling" autoscaling_group_name = aws_autoscaling_group.my_asg.name adjustment_type = "ChangeInCapacity" scaling_adjustment = "1" # add one instance cooldown = "300" # cooldown period after scaling } # ASG缩容策略 resource "aws_autoscaling_policy" "scale_down" { name = "asg-scale-down" autoscaling_group_name = aws_autoscaling_group.my_asg.name adjustment_type = "ChangeInCapacity" scaling_adjustment = "-1" cooldown = "300" policy_type = "SimpleScaling" }
1. 健康检查阈值过高
目标组健康检查的healthy_threshold和unhealthy_threshold均设为10,意味着新实例需要连续10次健康检查通过才能被标记为健康,而300秒的健康检查宽限期内很难完成足够次数的检查,直接导致ASG判定实例不健康并终止。
- 修复:将
healthy_threshold调整为2-3,unhealthy_threshold调整为3-5,降低健康检查的触发门槛。
2. 实例主动关机问题
错误明确显示Instance initiated shutdown,说明实例是自身触发关机,而非ASG强制操作,需排查:
- 检查实例用户数据、系统启动脚本(如systemd服务)是否存在错误的关机逻辑;
- 查看系统日志(
/var/log/messages、/var/log/syslog等),确认是否因资源耗尽、服务启动失败或硬件异常导致自动关机; - 验证自定义AMI的完整性,排除AMI本身存在初始化后自动关机的问题。
3. ASG配置矛盾
ASG代码中max_size、min_size、desired_capacity均设为1,但实际存在两台EC2实例,说明配置存在不一致:
- 若需保留两台实例,将ASG的
max_size、min_size、desired_capacity统一调整为2; - 检查AWS控制台中ASG的实际配置,确认是否与Terraform代码同步。
4. 健康检查宽限期不足
300秒宽限期对于启动较慢的服务(如需加载大量依赖、初始化外部连接的应用)可能不够,导致实例在完成初始化前就被判定为不健康。
- 修复:将
health_check_grace_period延长至600秒,给实例足够的启动时间。
内容的提问来源于stack exchange,提问作者DevNoob

