Terraform部署AWS架构遇ALB 502错误,EC2目标组健康检查失败
问题描述
使用Terraform在AWS上部署以下架构:
- 私有子网中部署3台自动扩缩容EC2实例(ASG),安装nginx
- 公有子网中部署ALB(应用负载均衡)
- 公有子网与私有子网位于不同可用区
- EC2从私有S3存储桶下载内容,通过nginx对外提供服务
- 限制ALB和服务器集群的入站访问仅允许80/TCP端口
当前访问ALB的DNS地址出现502 Bad Gateway错误,目标组内的EC2实例健康检查失败。提供的main.tf代码如下:
provider "aws" { region = "ap-southeast-1" } data "aws_vpc" "default" { default = true } data "aws_subnets" "public" { filter { name = "vpc-id" values = [data.aws_vpc.default.id] } } resource "aws_subnet" "private_subnet" { vpc_id = data.aws_vpc.default.id cidr_block = "172.31.32.0/20" availability_zone = "ap-southeast-1b" } # Elastic IP for NAT Gateway resource "aws_eip" "nat_eip" { domain = "vpc" } # NAT Gateway in the public subnet resource "aws_nat_gateway" "nat_gateway" { allocation_id = aws_eip.nat_eip.id subnet_id = data.aws_subnets.public.id } # Routing table for the private subnet resource "aws_route_table" "private_route_table" { vpc_id = data.aws_vpc.default.id route { cidr_block = "0.0.0.0/0" gateway_id = aws_nat_gateway.nat_gateway.id } } # Associate the routing table with the private subnet resource "aws_route_table_association" "private_route_association" { subnet_id = aws_subnet.private_subnet.id route_table_id = aws_route_table.private_route_table.id } resource "aws_security_group" "alb_sg" { vpc_id = data.aws_vpc.default.id ingress { from_port = 80 to_port = 80 protocol = "tcp" cidr_blocks = ["0.0.0.0/0"] } egress { from_port = 0 to_port = 0 protocol = "-1" cidr_blocks = ["0.0.0.0/0"] } } resource "aws_security_group" "ec2_sg" { vpc_id = data.aws_vpc.default.id ingress { from_port = 80 to_port = 80 protocol = "tcp" security_groups = [aws_security_group.alb_sg.id] } egress { from_port = 0 to_port = 0 protocol = "-1" cidr_blocks = ["0.0.0.0/0"] } } # IAM Role for EC2 to access S3 resource "aws_iam_role" "ec2_s3_role" { name = "EC2_S3_Access" assume_role_policy = jsonencode({ Version = "2012-10-17" Statement = [{ Effect = "Allow" Principal = { Service = "ec2.amazonaws.com" } Action = "sts:AssumeRole" }] }) } resource "aws_iam_policy" "s3_read_policy" { name = "S3ReadAccess" description = "Allows EC2 to read from S3 bucket" policy = jsonencode({ Version = "2012-10-17" Statement = [ { Effect = "Allow" Action = "s3:GetObject" Resource = "arn:aws:s3:::cherong-bucket/*" } ] }) } resource "aws_iam_role_policy_attachment" "ec2_s3_policy" { role = aws_iam_role.ec2_s3_role.name policy_arn = aws_iam_policy.s3_read_policy.arn } resource "aws_iam_instance_profile" "ec2_s3_profile" { name = "ec2-s3-profile" role = aws_iam_role.ec2_s3_role.name } # Launch Template resource "aws_launch_template" "sever_fleet_a" { name = "server-fleet-a" image_id = "ami-0599cde8e4a7ca305" instance_type = "t2.micro" iam_instance_profile { name = aws_iam_instance_profile.ec2_s3_profile.name } network_interfaces { associate_public_ip_address = false security_groups = [aws_security_group.ec2_sg.id] } user_data = base64encode(<<EOF #!/bin/bash sudo yum update -y sudo yum install -y nginx awscli aws s3 cp --recursive s3://cherong-bucket/webapp/index.html /usr/share/nginx/html/index.html sudo systemctl start nginx sudo systemctl enable nginx EOF ) } resource "aws_autoscaling_group" "asg" { vpc_zone_identifier = [aws_subnet.private_subnet.id] desired_capacity = 3 min_size = 3 max_size = 3 health_check_type = "ELB" launch_template { id = aws_launch_template.sever_fleet_a.id version = "$Latest" } } resource "aws_lb" "alb" { name = "public-alb" internal = false load_balancer_type = "application" security_groups = [aws_security_group.alb_sg.id] subnets = data.aws_subnets.public.ids } resource "aws_lb_target_group" "tg" { name = "tg-server-fleet-a" port = 80 protocol = "HTTP" vpc_id = data.aws_vpc.default.id health_check { path = "/" matcher = "200" healthy_threshold = 3 unhealthy_threshold = 3 timeout = 5 interval = 30 } } resource "aws_lb_listener" "listener" { load_balancer_arn = aws_lb.alb.arn port = 80 protocol = "HTTP" default_action { type = "forward" target_group_arn = aws_lb_target_group.tg.arn } } resource "aws_autoscaling_attachment" "asg_attachment" { autoscaling_group_name = aws_autoscaling_group.asg.id lb_target_group_arn = aws_lb_target_group.tg.arn } output "alb_dns_name" { value = aws_lb.alb.dns_name description = " The domain name of the load balancer" }
问题排查与修复
首先明确:ALB和EC2的安全组配置本身没有问题——EC2安全组允许ALB安全组访问80端口,ALB安全组允许全网访问80,符合通信要求。502错误和健康检查失败的根源在其他配置项:
1. NAT网关子网配置错误
aws_nat_gateway资源中,subnet_id = data.aws_subnets.public.id是无效的,因为data.aws_subnets.public返回的是子网集合,没有id属性,需指定具体的公有子网:
resource "aws_nat_gateway" "nat_gateway" { allocation_id = aws_eip.nat_eip.id subnet_id = data.aws_subnets.public.ids[0] # 取第一个公有子网 }
2. 用户数据的S3拷贝命令错误
aws s3 cp --recursive s3://cherong-bucket/webapp/index.html ...中,--recursive用于同步目录,但目标是单个文件,会导致命令执行失败,nginx启动后无有效首页文件,健康检查失败。修改为:
- 若仅拷贝单个index.html:
aws s3 cp s3://cherong-bucket/webapp/index.html /usr/share/nginx/html/index.html
- 若需同步整个webapp目录:
aws s3 cp --recursive s3://cherong-bucket/webapp/ /usr/share/nginx/html/
3. ASG健康检查 grace period 缺失
实例初始化需要时间完成nginx安装、S3文件下载,默认的健康检查判定逻辑可能过早标记实例不健康。添加health_check_grace_period给实例足够的初始化时间:
resource "aws_autoscaling_group" "asg" { vpc_zone_identifier = [aws_subnet.private_subnet.id] desired_capacity = 3 min_size = 3 max_size = 3 health_check_type = "ELB" health_check_grace_period = 120 # 预留2分钟初始化时间 launch_template { id = aws_launch_template.sever_fleet_a.id version = "$Latest" } }
4. 额外验证项
若修复后仍有问题,需检查:
- 查看EC2实例的系统日志,确认nginx是否正常启动,
systemctl status nginx是否返回active状态 - 确认S3文件拷贝成功,
/usr/share/nginx/html/index.html文件存在且内容正常 - 验证私有子网的路由表是否正确关联NAT网关,确保实例能访问外网下载nginx和S3资源
内容的提问来源于stack exchange,提问作者chessy
相关产品推荐
相关产品推荐

