You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ansible wait_for_connection失效?GCP创建VM后自动等待连接求助

Got it, let's figure out why your wait_for_connection isn't working as expected and fix this once and for all.

First, the core issue here is likely Ansible's default behavior of auto-gathering facts before running any tasks—if it can't connect to gather facts, it fails immediately, skipping your wait_for_connection task entirely. Plus, the default parameters for wait_for_connection might not be tuned to GCP's VM boot timeline.

Here's the fixed playbook that addresses both problems:

- hosts: all
  gather_facts: no  # Disable auto-fact gathering to avoid early connection failure
  tasks:
    - name: Wait for SSH service to be fully available
      wait_for_connection:
        delay: 10  # Wait 10s first before starting checks (gives VM time to initialize network)
        timeout: 300  # Allow up to 5 minutes for the VM to get ready
        sleep: 5  # Retry connection every 5 seconds

    - name: Manually gather facts now that connection works
      setup:

    - name: Verify connectivity with ping
      ping:

Why this works:

  1. gather_facts: no: By turning off automatic fact gathering, we prevent Ansible from trying to connect to the VM before our wait task runs. This is the most common reason the initial setup fails.
  2. Tuned wait_for_connection parameters:
    • delay:10 gives GCP a moment to assign the external IP and bring up the network interface before we start probing.
    • timeout:300 ensures we don't give up too early (GCP VMs can take a minute or two to fully boot SSH, especially with custom images).
    • sleep:5 balances between checking frequently enough and not overwhelming the VM with connection attempts.

Bonus: Even more reliability for GCP

If you want to ensure the VM is in a RUNNING state at the GCP level before checking SSH, you can add a pre-check using GCP's Ansible modules (requires google-cloud-sdk and ansible-collection-google installed):

# First, verify the GCP instance status from your control node
- hosts: localhost
  gather_facts: no
  vars:
    gcp_project: "your-gcp-project-id"
    gcp_zone: "your-vm-zone"
    vm_name: "your-target-vm-name"
  tasks:
    - name: Fetch GCP instance details
      gcp_compute_instance_info:
        project: "{{ gcp_project }}"
        zone: "{{ gcp_zone }}"
        filters:
          - name = "{{ vm_name }}"
      register: vm_details

    - name: Wait for VM to reach RUNNING state
      wait_for:
        timeout: 300
      until: vm_details.resources[0].status == "RUNNING"
      retries: 30
      delay: 10

    - name: Add the VM's external IP to inventory
      add_host:
        name: "{{ vm_details.resources[0].networkInterfaces[0].accessConfigs[0].natIP }}"
        groups: gcp_target_vms

# Now connect to the VM and wait for SSH
- hosts: gcp_target_vms
  gather_facts: no
  tasks:
    - name: Wait for SSH connectivity
      wait_for_connection:
        delay: 5
        timeout: 300
        sleep: 3

    - name: Ping test to confirm connectivity
      ping:

Quick sanity checks to rule out other issues:

  • Ensure your GCP firewall allows inbound SSH (port 22) from your Ansible control node's IP.
  • Verify the SSH user in your Ansible inventory matches the default user for your GCP VM image (e.g., ubuntu for Ubuntu images, debian for Debian).

内容的提问来源于stack exchange,提问作者Jananath Banuka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 16:53:13