Ansible中使用block模块实现AWS EC2实例错误回滚失败求助
Hey there! Let's break down why your block/rescue setup isn't working and fix it step by step. The main issues are around how you're using handlers, task structure, and variable references. Here's what's going wrong and how to fix it:
1. You're using notify incorrectly in the block
notify only triggers handlers after successful tasks, and tucking your instance creation into a handler means it won't run the way you expect. Handlers are meant for deferred, post-success actions (like restarting a service), not core tasks like spinning up EC2 instances. You should run the instance creation task directly inside the block, not via a debug task's notification.
2. Rescue can't rely on new_ec2 if the instance creation failed
If any task in the block fails before the instance is fully created, the new_ec2 variable might not exist at all. Referencing new_ec2.instances[0].id directly in the rescue will throw an error—we need to add checks to make sure the variable exists before trying to clean up.
3. Variable reference mistakes
- You used
{{ region }}instead of{{ aws_vars.region }}in the EC2 creation task (this would throw an undefined variable error) new_ec2.idis invalid; the instance ID lives undernew_ec2.instances[0].id(or.instance_id—both work, just be consistent)- The
with_items: "{{ new_ec2 }}"in the CloudWatch task is unnecessary since you're only creating one instance.
4. Handlers aren't the right tool for this flow
Handlers are not designed to run critical setup tasks like instance creation. Let's move those tasks directly into the block where they belong.
Here's the corrected Playbook that implements proper error rollback, including your original volume attachment task:
--- # EC2 Migrations. - hosts: localhost connection: local gather_facts: no tasks: - name: Load AWS variables include_vars: dir: files name: aws_vars - name: Create EC2 instance and related resources block: - name: Launch EC2 instance ec2: key_name: "{{ aws_vars.key_name }}" group: "{{ aws_vars.security_group }}" instance_type: "{{ aws_vars.instance_type }}" image: "{{ aws_vars.image }}" region: "{{ aws_vars.region }}" wait: yes count: 1 monitoring: yes assign_public_ip: yes register: new_ec2 - debug: msg: "Created EC2 instance with ID: {{ new_ec2.instances[0].id }}" - name: Attach existing volume ec2_vol: device_name: xvdf instance: "{{ new_ec2.instances[0].id }}" region: "{{ aws_vars.region }}" id: "{{ volume_id }}" # Ensure this variable is defined in your vars files delete_on_termination: yes - name: Create CloudWatch CPU alarm ec2_metric_alarm: state: present region: "{{ aws_vars.region }}" name: "{{ new_ec2.instances[0].id }}-High-CPU" metric: "CPUUtilization" namespace: "AWS/EC2" statistic: Average comparison: ">=" threshold: "90.0" period: 300 evaluation_periods: 3 unit: "Percent" description: "Instance CPU is above 90%" dimensions: InstanceId: "{{ new_ec2.instances[0].id }}" alarm_actions: "{{ aws_vars.sns_arn }}" ok_actions: "{{ aws_vars.sns_arn }}" rescue: - name: Rollback - Terminate invalid EC2 instance (if created) ec2: instance_ids: "{{ new_ec2.instances[0].id | default('') }}" region: "{{ aws_vars.region }}" state: terminated when: new_ec2 is defined and new_ec2.instances | length > 0 ignore_errors: yes # Make sure rollback doesn't fail if the instance never launched - debug: msg: "Rollback completed: Terminated any partially created EC2 resources"
Key Fixes Explained:
- Direct task execution in block: All critical setup tasks (instance launch, volume attachment, CloudWatch alarm) live directly in the
block—if any fail, therescuesection triggers immediately. - Safe rollback logic: The rescue task checks if
new_ec2exists and has instances before attempting termination, and usesignore_errors: yesto avoid breaking the rollback if the instance wasn't created at all. - Fixed variable references: Corrected all undefined variable issues and cleaned up invalid instance ID references.
- Terminate instead of stop: I switched
state: stoppedtostate: terminatedin the rollback since you wanted to clean up invalid instances. If you prefer to stop instead, just change that value back.
内容的提问来源于stack exchange,提问作者Tomer Leibovich

