You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Terraform销毁aws_ecs_cluster时local-exec报aws命令未找到问题问询

Terraform销毁AWS ECS集群报错解决方案

问题1:local-exec执行时报 /bin/bash: aws: command not found 错误

该报错原因是Terraform执行local-exec置备器时,运行环境的PATH变量未包含AWS CLI的安装路径,可通过两种方式修复:

  • 方式一:指定AWS CLI绝对路径
    本地执行which aws获取AWS CLI的实际安装路径(常见路径为/usr/local/bin/aws),将脚本中所有aws命令替换为绝对路径即可。
  • 方式二:脚本开头导入PATH变量
    在local-exec的command代码块第一行新增export PATH=$PATH:/usr/local/bin:/usr/bin,覆盖绝大多数场景下AWS CLI的安装路径。

此外需要确认执行terraform destroy的用户,已经配置了拥有ECS、Auto Scaling Group操作权限的AWS凭证,凭证规则与AWS CLI通用(支持环境变量、~/.aws/credentials配置文件、IAM角色等方式)。


问题2:销毁集群抛出ClusterContainsContainerInstancesException错误、偶现IGW分离超时

该问题是因为原有置备器只调整了ASG容量,没有等待实例完全销毁就触发了ECS集群、VPC资源的删除逻辑,修复方式如下:

第一步:修改置备器逻辑,新增实例销毁等待

修改后的完整aws_ecs_cluster资源代码如下:

resource "aws_ecs_cluster" "demo" {
  name = var.ecs_cluster_name

  capacity_providers = [local.cluster_name]

  default_capacity_provider_strategy {
    capacity_provider = local.cluster_name
  }

  # 销毁集群前先终止所有实例
  provisioner "local-exec" {
    when = destroy
    interpreter = ["bash", "-c"]
    command = <<EOT
      # 导入PATH解决aws命令找不到问题
      export PATH=$PATH:/usr/local/bin:/usr/bin
      set -e
      CAP_PROVS="$(aws ecs describe-clusters --clusters "${self.arn}" \
        --query 'clusters[*].capacityProviders[*]' --output text)"

      ASG_ARNS="$(aws ecs describe-capacity-providers \
        --capacity-providers "$CAP_PROVS" \
        --query 'capacityProviders[*].autoScalingGroupProvider.autoScalingGroupArn' \
        --output text)"

      if [ -n "$ASG_ARNS" ] && [ "$ASG_ARNS" != "None" ]
      then
        for ASG_ARN in $ASG_ARNS
        do
          ASG_NAME=$(echo $ASG_ARN | cut -d/ -f2-)

          aws autoscaling update-auto-scaling-group \
            --auto-scaling-group-name "$ASG_NAME" \
            --min-size 0 --max-size 0 --desired-capacity 0

          INSTANCES="$(aws autoscaling describe-auto-scaling-groups \
            --auto-scaling-group-names "$ASG_NAME" \
            --query 'AutoScalingGroups[*].Instances[*].InstanceId' \
            --output text)"
          if [ -n "$INSTANCES" ] && [ "$INSTANCES" != "None" ]; then
            aws autoscaling set-instance-protection --instance-ids $INSTANCES \
              --auto-scaling-group-name "$ASG_NAME" \
              --no-protected-from-scale-in
            # 等待ASG内所有实例完全销毁
            aws autoscaling wait no-instances-running --auto-scaling-group-names "$ASG_NAME"
          fi
        done
      fi
      # 等待集群状态变为无活跃实例
      aws ecs wait cluster-inactive --cluster "${self.arn}"
    EOT
  }
}

第二步:新增销毁延迟避免IGW分离超时

如果仍出现互联网网关分离超时,可新增等待资源,让VPC相关资源延迟销毁:

# 销毁等待资源
resource "time_sleep" "destroy_wait" {
  create_duration = "0s"
  destroy_duration = "60s"

  triggers = {
    ecs_cluster_arn = aws_ecs_cluster.demo.arn
  }

  lifecycle {
    create_before_destroy = true
  }
}

# 给IGW资源添加依赖,示例如下,原有配置保持不变
resource "aws_internet_gateway" "demo" {
  vpc_id = aws_vpc.demo.id
  # 其他配置...

  depends_on = [time_sleep.destroy_wait]
}

可选优化方案(更符合Terraform声明式逻辑)

如果不想使用过程式的local-exec置备器,可以调整资源依赖顺序:将容量提供者、ASG资源的依赖设置为依赖ECS集群,这样Terraform销毁时会自动先删除ASG,再删除ECS集群,从根源上避免实例残留问题。

内容的提问来源于stack exchange,提问作者asdzxsad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 04:15:04