如何在Chef应用版本变更时触发Terraform重建AWS实例?
Let's break down why your current setup isn't working first: when you update the app version in some_service_user_data.sh, Terraform will create a new launch configuration (thanks to your create_before_destroy lifecycle rule), but the existing Auto Scaling Group (ASG) won't automatically replace its running instances. New instances launched by the ASG (e.g., during scaling events) will use the new LC, but your existing ones stay on the old version.
Here are three solid solutions to get the full instance replacement you want:
1. Use a Versioned Tag + ASG Lifecycle Rule
This approach ties your app version directly to the ASG's configuration, ensuring Terraform triggers a full ASG recreate when the version changes.
First, define a Terraform variable for your app version (to avoid duplicating values):
variable "some_service_version" { type = string default = "18.01.124-v02" }
Update your launch configuration to use a template for user data (instead of a static file) so it pulls the version from the variable:
resource "aws_launch_configuration" "some_service" { image_id = lookup(var.aws_amis, var.aws_region) instance_type = var.instance_type security_groups = [aws_security_group.some_service.id] key_name = var.key_name # Use templatefile to inject the version variable user_data = templatefile("sh/some_service_user_data.sh", { service_version = var.some_service_version }) lifecycle { create_before_destroy = true } }
Modify some_service_user_data.sh to use the injected variable:
#!/bin/bash -xev cd /etc/chef/ # Install chef curl -L https://omnitruck.chef.io/install.sh | bash || error_exit 'could not install chef' # Create first-boot.json with injected version cat > "/etc/chef/first-boot.json" << EOF { "some_service": { "environment": "aws", "version": "${service_version}" }, "run_list" :[ "role[some_service]" ] } EOF NODE_NAME=`hostname` # Create client.rb cat > "/etc/chef/client.rb" << EOF log_level :info log_location STDOUT chef_server_url 'https://chef-server/organizations/myorg' validation_client_name 'myorg-validator' validation_key '/etc/chef/myorg-validator.pem' node_name "${NODE_NAME}" ssl_verify_mode :verify_none EOF sudo chef-client -j /etc/chef/first-boot.json -E 'aws' service some_service start
Finally, add a versioned tag to your ASG and enable create_before_destroy to force a full recreate when the version changes:
resource "aws_autoscaling_group" "some_service" { launch_configuration = aws_launch_configuration.some_service.id availability_zones = split(",", var.availability_zones) depends_on = [aws_instance.some_other_resource] min_size = 1 max_size = 3 tag { key = "Name" value = "terraform_asg_some_service" propagate_at_launch = true } # Add a version tag that propagates to instances tag { key = "ServiceVersion" value = var.some_service_version propagate_at_launch = true } lifecycle { create_before_destroy = true } }
Now, whenever you update var.some_service_version and run terraform apply, Terraform will:
- Create a new launch configuration with the updated user data
- Spin up a new ASG with the version tag
- Destroy the old ASG (and its instances) once the new one is healthy
2. Use Instance Refresh with replace_triggered_by (Terraform 0.13+)
This is the most native Terraform approach—no extra tags needed. It leverages the ASG's instance refresh feature to replace running instances when the launch configuration changes.
Update your ASG configuration like this:
resource "aws_autoscaling_group" "some_service" { launch_configuration = aws_launch_configuration.some_service.id availability_zones = split(",", var.availability_zones) depends_on = [aws_instance.some_other_resource] min_size = 1 max_size = 3 tag { key = "Name" value = "terraform_asg_some_service" propagate_at_launch = true } # Trigger instance refresh when the launch configuration changes replace_triggered_by = [aws_launch_configuration.some_service.id] # Configure a rolling refresh to avoid downtime instance_refresh { strategy = "Rolling" preferences { min_healthy_percentage = 50 # Adjust based on your availability needs } } lifecycle { create_before_destroy = true } }
When you update your app version (and thus the user data/launch configuration), Terraform will detect the LC change and trigger an instance refresh. The ASG will gradually replace old instances with new ones using the updated LC, maintaining your desired availability percentage.
3. Force Instance Replacement via AWS CLI (Legacy Approach)
If you're using an older Terraform version that doesn't support replace_triggered_by, you can use a null_resource to run an AWS CLI command that starts an instance refresh:
resource "null_resource" "refresh_asg_instances" { # Trigger this whenever the launch configuration changes triggers = { launch_config_id = aws_launch_configuration.some_service.id } provisioner "local-exec" { command = "aws autoscaling start-instance-refresh --auto-scaling-group-name ${aws_autoscaling_group.some_service.name} --strategy Rolling --preferences MinHealthyPercentage=50" } }
Note: This requires the AWS CLI to be installed and configured on the machine running terraform apply.
My recommendation is Solution 2—it's clean, uses Terraform's native features, and avoids unnecessary ASG recreates by using rolling instance refreshes instead of full ASG replacement.
内容的提问来源于stack exchange,提问作者gbaii

