Terraform部署AWS EC2实例后无代码变更却触发销毁重建的原因及预防方案
Terraform检测到EC2实例外部变更并计划重建的原因及预防方案
问题背景
我用这段Terraform代码创建了AWS EC2实例:
resource "aws_instance" "example" { ami = var.ami-id instance_type = var.ec2_type key_name = var.keyname subnet_id = "subnet-05a63e5c1a6bcb7ac" security_groups = ["sg-082d39ed218fc0f2e"] # root disk root_block_device { volume_size = "10" volume_type = "gp3" encrypted = true delete_on_termination = true } tags = { Name = var.instance_name Environment = "dev" } metadata_options { http_endpoint = "enabled" http_put_response_hop_limit = 1 http_tokens = "required" } }
执行terraform apply完成5分钟后,我没改任何代码,运行terraform plan却发现Terraform检测到外部变更,要销毁重建这个EC2实例。terraform plan的输出如下:
aws_instance.example: Refreshing state... [id=i-0aa279957d1287100] Note: Objects have changed outside of Terraform Terraform detected the following changes made outside of Terraform since the last "terraform apply": # aws_instance.example has been changed ~ resource "aws_instance" "example" { id = "i-0aa279957d1287100" ~ security_groups = [ - "sg-082d39ed218fc0f2e", ] tags = { "Environment" = "dev" "Name" = "ec2linux" } # (26 unchanged attributes hidden) ~ root_block_device { + tags = {} # (9 unchanged attributes hidden) } # (4 unchanged blocks hidden) } Unless you have made equivalent changes to your configuration, or ignored the relevant attributes using ignore_changes, the following plan may include actions to undo or respond to these changes. ───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── Terraform used the selected providers to generate the following execution plan. Resource actions are indicated with the following symbols: -/+ destroy and then create replacement
原因分析
咱们先从terraform plan的输出里抓关键信息:
- 核心触发点:安全组被移除
原本配置里的sg-082d39ed218fc0f2e安全组被从实例上删掉了。对于EC2实例来说,security_groups是Terraform需要严格维护的属性,一旦实际状态和代码配置不一致,而且这个属性的修改无法通过在线更新完成(部分EC2属性必须重建实例才能生效),Terraform就会计划销毁旧实例、创建新实例来匹配配置。 - 次要变更:根磁盘空标签
根磁盘多了个空的tags = {},这个一般不会触发重建,只是顺带检测到的外部变更,真正导致重建的还是安全组的变化。
至于安全组为啥会被移除,常见的几种情况:
- 有人手动在AWS控制台/CLI里把这个安全组从实例上解绑了
- 其他自动化工具(比如AWS Config规则、自定义运维脚本、甚至其他IaC工具)修改了实例的安全组配置
- 极端情况:如果安全组和实例不在同一个VPC,AWS会自动清理非法关联,但这种情况在
terraform apply的时候就会报错,所以概率很低
预防方案
针对这类外部变更导致的资源重建问题,咱们可以从这几个维度入手解决:
1. 杜绝非Terraform的资源修改
- 给Terraform管理的所有资源统一加个标签,比如
ManagedBy = "Terraform",然后通过AWS IAM策略限制:只有Terraform使用的角色能修改带这个标签的资源,其他用户/角色只能查看。 - 团队内部定好规矩:所有基础设施变更必须走Terraform代码提交,禁止直接在控制台随意修改。
2. 用ignore_changes忽略特定属性
如果某些属性确实需要被外部工具修改,而且你不想让Terraform干预这些变化,可以在资源块里加lifecycle配置忽略这些属性:
resource "aws_instance" "example" { # 原有代码... lifecycle { ignore_changes = [ security_groups, # 忽略安全组的外部变更 root_block_device[0].tags, # 忽略根磁盘标签的变更 ] } }
这里要注意:用ignore_changes得谨慎,确认这些属性的一致性不需要Terraform维护才行,不然容易出现配置和实际状态脱节的问题。
3. 不要硬编码安全组ID,用Terraform资源引用
你现在代码里硬编码了安全组ID,推荐改成用Terraform自己管理安全组,然后通过资源引用关联:
# 先定义安全组 resource "aws_security_group" "example" { name = "example-ec2-sg" description = "Security group for example EC2 instance" vpc_id = var.vpc_id # 和实例的子网属于同一个VPC # 根据需求添加入站/出站规则 ingress { from_port = 22 to_port = 22 protocol = "tcp" cidr_blocks = ["192.168.0.0/24"] # 实际生产环境请限制定义的IP段 } egress { from_port = 0 to_port = 0 protocol = "-1" cidr_blocks = ["0.0.0.0/0"] } } # 然后在实例里引用这个安全组 resource "aws_instance" "example" { # 原有代码... # 如果是VPC环境,更推荐用vpc_security_group_ids而不是security_groups vpc_security_group_ids = [aws_security_group.example.id] }
这样安全组的生命周期也由Terraform管理,就不会出现外部随便修改的情况了。
4. 启用Terraform状态锁定
用Terraform Cloud或者AWS S3+DynamoDB来做状态存储和锁定,确保同一时间只有一个人/工具能修改Terraform状态,避免并发操作导致的状态不一致问题。
5. 定期检查状态一致性
可以把terraform plan加入CI/CD流程,定期跑一下,及时发现外部变更并处理,避免问题积累到需要重建资源的地步。
内容的提问来源于stack exchange,提问作者sfgroups
相关产品推荐
相关产品推荐

