Terraform配置EC2容量提供商时ECS任务持续处于PROVISIONING状态
ECS任务卡在PROVISIONING状态,实例未正确注册到集群(使用Capacity Providers + Terraform)
我通过Terraform配置了ECS Capacity Providers,期望调整服务期望计数时自动扩缩容EC2实例。当前问题:调整服务期望计数后,Auto Scaling告警正常触发,实例扩容操作执行成功,但服务创建的新任务始终停留在PROVISIONING状态,无法转为RUNNING。进一步排查发现,扩容的EC2实例并未正确注册到ECS集群,已排除user_data配置错误导致实例注册到默认集群的情况。
相关Terraform配置
本地变量与Launch Template
locals { ami_id = jsondecode(data.aws_ssm_parameter.ecs_optimized_ami.value)["image_id"] cluster_user_data = "${base64encode(<<EOF #! /bin/bash sudo apt-get update sudo echo "ECS_CLUSTER=${aws_ecs_cluster.cluster.name}" >> /etc/ecs/ecs.config EOF )}" } resource "aws_launch_template" "ecs-public" { name_prefix = "ecs-${var.cluster_name}-public" image_id = local.ami_id instance_type = "t3.small" iam_instance_profile { name = aws_iam_instance_profile.cluster.name } block_device_mappings { device_name = "/dev/xvda" ebs { volume_size = var.cluster_instance_root_block_device_size volume_type = var.cluster_instance_root_block_device_type } } # TODO use Dynamic here network_interfaces { device_index = 0 security_groups = var.security_groups_external delete_on_termination = true subnet_id = var.public_subnet_ids[0] } network_interfaces { device_index = 1 security_groups = var.security_groups_external delete_on_termination = true subnet_id = var.public_subnet_ids[1] } user_data = local.cluster_user_data key_name = aws_key_pair.generated_key[0].key_name lifecycle { create_before_destroy = true } depends_on = [ null_resource.iam_wait ] }
Auto Scaling Group
resource "aws_autoscaling_group" "cluster_public" { name_prefix = "asg-public-${var.cluster_name}" vpc_zone_identifier = var.public_subnet_ids launch_template { id = aws_launch_template.ecs-public.id version = "$Latest" } min_size = 1 max_size = 5 desired_capacity = 1 protect_from_scale_in = false tag { key = "Name" value = "worker-public-${var.cluster_name}" propagate_at_launch = true } tag { key = "ClusterName" value = var.cluster_name propagate_at_launch = true } tag { key = "AmazonECSManaged" value = true propagate_at_launch = true } dynamic "tag" { for_each = var.tags content { key = tag.key value = tag.value propagate_at_launch = true } } lifecycle { create_before_destroy = true } }
Capacity Provider与集群关联
resource "aws_ecs_capacity_provider" "autoscaling_group_public" { name = "cp-${var.cluster_name}-public" auto_scaling_group_provider { auto_scaling_group_arn = aws_autoscaling_group.cluster_public.arn managed_termination_protection = "DISABLED" managed_scaling { status = "ENABLED" target_capacity = 100 minimum_scaling_step_size = 1 maximum_scaling_step_size = 100 } } } resource "aws_ecs_cluster_capacity_providers" "cluster_capacity_providers" { cluster_name = aws_ecs_cluster.cluster.name capacity_providers = [aws_ecs_capacity_provider.autoscaling_group_private[0].name, aws_ecs_capacity_provider.autoscaling_group_public[0].name] }
服务与任务模块配置
module "container_definition" { source = "cloudposse/ecs-container-definition/aws" version = "0.58.1" container_name = local.container_name container_image = "${module.global_settings.aws_account_id}.dkr.ecr.${module.global_settings.region}.amazonaws.com/${local.project_name}:Staging-latest" container_memory = 512 container_memory_reservation = 256 container_cpu = 256 essential = true readonly_root_filesystem = false environment = local.task_environment_variables port_mappings = local.port_mappings log_configuration = local.container_log_configuration } module "ecs_alb_service_task" { source = "cloudposse/ecs-alb-service-task/aws" version = "0.66.2" namespace = var.cluster_name stage = "Staging" name = local.project_name attributes = [] container_definition_json = module.container_definition.sensitive_json_map_encoded_list #Load Balancer alb_security_group = var.security_group_id ecs_load_balancers = local.ecs_load_balancer_config #VPC vpc_id = var.vpc_id subnet_ids = var.subnet_ids network_mode = "awsvpc" #Capacity Provider Strategy capacity_provider_strategies = [ { capacity_provider = var.capacity_provider_name weight = 1 base = 0 } ] desired_count = 2 launch_type = "EC2" ignore_changes_desired_count = true ecs_cluster_arn = var.cluster_arn security_group_ids = [var.security_group_id] ignore_changes_task_definition = true health_check_grace_period_seconds = 200 deployment_minimum_healthy_percent = 100 deployment_maximum_percent = 200 deployment_controller_type = "ECS" task_memory = 512 task_cpu = 256 force_new_deployment = true ordered_placement_strategy = [ { type = "spread" field = "attribute:ecs.availability-zone" },{ type = "spread" field = "instanceId" } ] label_order = local.label_order labels_as_tags = local.labels_as_tags propagate_tags = local.propagate_tags tags = merge(var.tags, local.tags) task_exec_role_arn = [module.task_excecution_role.task_excecution_role_arn] task_role_arn = [module.task_excecution_role.task_excecution_role_arn] }
已尝试的排查操作
- 多次调整服务期望计数
- 删除并重新创建ECS服务
- 设置
ignore_changes_desired_count = true - 调整
capacity_provider_strategies的base=1、weight=3 - 将实例类型从
t3.micro改为t3.small
内容的提问来源于stack exchange,提问作者Math.Random
相关产品推荐
相关产品推荐

