You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

EC2实例结合ALB与ASG异常终止重建问题排查求助

问题描述

我搭建了两台EC2服务器,关联到ASG(自动伸缩组)并挂载到ALB(应用负载均衡),分别位于eu-west-2a和eu-west-2b可用区。这些新服务器持续被终止并重建,无法定位原因。


相关配置

ALB目标组配置

resource "aws_lb_target_group" "targetgroup" {
  name         = "targetgroup"
  target_type  = "instance"
  port         = 80
  protocol     = "HTTP"
  vpc_id       = aws_vpc.main_vpc.id

  health_check {
    path                = "/"
    port                = 80
    protocol            = "HTTP"
    healthy_threshold   = 10
    unhealthy_threshold = 10
    matcher             = "200-499"
  }
}

安全组已开放80端口。

ASG活动错误信息

  • 其中一条错误:

Terminating EC2 instance: i-04dbb4d0f8c06b355 - Waiting For ELB Connection Draining. At 2024-03-11T15:02:16Z an instance was taken out of service in response to an EC2 health check indicating it has been terminated or stopped.

  • 更常见的错误:

Launching a new EC2 instance: i-0d4840406c9612b45. Status Reason: Instance became unhealthy while waiting for instance to be in InService state. Termination Reason: Client.InstanceInitiatedShutdown: Instance initiated shutdown

ASG配置

resource "aws_autoscaling_group" "my_asg" {
  name                      = "my_asg"
  max_size                  = 1
  min_size                  = 1
  health_check_type         = "ELB"    # optional
  desired_capacity          = 1
  target_group_arns = [aws_lb_target_group.targetgroup.arn]
  health_check_grace_period  = 300 
  vpc_zone_identifier       = [aws_subnet.public-subnet-1.id, aws_subnet.public-subnet-2.id]
}

伸缩策略配置

# ASG扩容策略
resource "aws_autoscaling_policy" "scale_up" {
  name                   = "scale_up"
  policy_type            = "SimpleScaling"
  autoscaling_group_name = aws_autoscaling_group.my_asg.name
  adjustment_type        = "ChangeInCapacity"
  scaling_adjustment     = "1"      # add one instance
  cooldown               = "300"    # cooldown period after scaling
}

# ASG缩容策略
resource "aws_autoscaling_policy" "scale_down" {
  name                   = "asg-scale-down"
  autoscaling_group_name = aws_autoscaling_group.my_asg.name
  adjustment_type        = "ChangeInCapacity"
  scaling_adjustment     = "-1"
  cooldown               = "300"
  policy_type            = "SimpleScaling"
}

排查与解决方案

1. 健康检查阈值过高

目标组健康检查的healthy_threshold和unhealthy_threshold均设为10,意味着新实例需要连续10次健康检查通过才能被标记为健康,而300秒的健康检查宽限期内很难完成足够次数的检查,直接导致ASG判定实例不健康并终止。

  • 修复:将healthy_threshold调整为2-3,unhealthy_threshold调整为3-5,降低健康检查的触发门槛。

2. 实例主动关机问题

错误明确显示Instance initiated shutdown,说明实例是自身触发关机,而非ASG强制操作,需排查:

  • 检查实例用户数据、系统启动脚本(如systemd服务)是否存在错误的关机逻辑;
  • 查看系统日志(/var/log/messages、/var/log/syslog等),确认是否因资源耗尽、服务启动失败或硬件异常导致自动关机;
  • 验证自定义AMI的完整性,排除AMI本身存在初始化后自动关机的问题。

3. ASG配置矛盾

ASG代码中max_size、min_size、desired_capacity均设为1,但实际存在两台EC2实例,说明配置存在不一致:

  • 若需保留两台实例,将ASG的max_size、min_size、desired_capacity统一调整为2;
  • 检查AWS控制台中ASG的实际配置,确认是否与Terraform代码同步。

4. 健康检查宽限期不足

300秒宽限期对于启动较慢的服务(如需加载大量依赖、初始化外部连接的应用)可能不够,导致实例在完成初始化前就被判定为不健康。

  • 修复:将health_check_grace_period延长至600秒,给实例足够的启动时间。

内容的提问来源于stack exchange,提问作者DevNoob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 01:49:53