You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Terraform从模板部署VM时自定义失败:需等待首次启动完成

基于Terraform+vSphere的VM自定义部署故障排查

环境信息

  • Terraform版本:v1.10
  • 目标平台:vCenter (vSphere) v7.x
  • vSphere插件版本:v2.10

VM模板详情

  • 操作系统:Ubuntu 24.04.1 LTS
  • 构建流程:Packer → cloud-config → autoinstall

Packer模板构建流程

Packer构建VM的最后步骤:

  • 第一次重启完成cloud-init配置
  • Packer验证VM可通过SSH访问

完成后VM会被关机,转换为模板或导入内容库。因此从该模板克隆的VM,从操作系统角度属于第二次启动。

模板克隆后的第二次启动逻辑

模板的user-data中通过runcmd创建了名为second-boot-service的systemd服务,作用如下:

  1. 验证系统启动次数
  2. 若为第二次启动,禁用该服务,执行cloud-init clean并重启系统

该逻辑确保克隆后的新VM具备以下特性:

  • 生成新的/etc/machine-id
  • 生成新的SSH密钥
  • 清理日志

注:手动部署且不做操作系统自定义时,该逻辑运行正常。

故障场景与现象

使用Terraform通过vsphere_virtual_machine.clone.customize.linux_options从模板部署VM时,流程如下:

  1. 执行克隆操作
  2. VM第一次启动
  3. Terraform尝试应用自定义配置

报错信息

vsphere_virtual_machine.hcv_nodes[0]: Still creating... [1m50s elapsed]
vsphere_virtual_machine.hcv_nodes[0]: Still creating... [2m0s elapsed]
╷
│ Error: 
│ Virtual machine customization failed on "/CDADC001/vm/cdavault001-01.domain.com":
│ 
│ An error occurred while customizing VM cdavault001-01.domain.com. For details reference the log file /var/log/vmware-imc/toolsDeployPkg.log in the guest OS.

故障原因

second-boot-service执行cloud-init clean命令后触发系统重启,但此时Terraform仍在尝试应用linux_options自定义配置,导致冲突。

需求

  1. 让镜像在更晚的时机执行cloud-init clean,且确保命令一定会运行
  2. 让Terraform在尝试自定义前等待系统完成初始化

相关代码片段

user-data(片段)

user-data:
        # Post-Processing
        # - Execute these on the 2nd boot, only once (poorman's rc.local)
        # logic: since Packer/VMware-ISO has already booted the system once to validate SSH connectivity,
        # this cleanup process needs to run the 1st time the final image boots.
        # This will reset SSH keys, machineID, and cloud-init logs.
        write_files:
            - path: /usr/local/bin/second-boot-script.sh
              permissions: '0755'
              owner: root:root
              content: |
                #!/bin/bash
                # Path to the counter file
                COUNTER_FILE="/var/run/boot-counter"

                # Systemd service name
                SERVICE_NAME="second-boot.service"

                # Initialize boot count if the file doesn't exist
                if [ ! -f "$COUNTER_FILE" ]; then
                    echo 1 > "$COUNTER_FILE"
                fi

                # Read the current boot count
                BOOT_COUNT=$(cat "$COUNTER_FILE")

                if [ "$BOOT_COUNT" -eq 1 ]; then
                    echo "Executing second boot commands..."

                    # Place all pre-reboot commands here
                    echo "Running pre-reboot commands..." >> /var/log/second-boot.log
                    # Example: Updating some configurations or running additional setup
                    echo "Pre-reboot tasks completed at $(date)" >> /var/log/second-boot.log

                    # Disable the systemd service to prevent future executions
                    echo "Disabling $SERVICE_NAME to prevent future runs."
                    systemctl disable "$SERVICE_NAME"

                    # Final step: Run cloud-init clean and reboot
                    echo "Running cloud-init clean and rebooting system..." >> /var/log/second-boot.log
                    cloud-init clean --logs --machine-id --configs ssh_config --reboot
                fi

                # Increment the boot count
                BOOT_COUNT=$((BOOT_COUNT + 1))
                echo "$BOOT_COUNT" > "$COUNTER_FILE"

        runcmd:
        - echo "Creating systemd service for second-boot-script.sh"
        - |
            cat <<EOF > /etc/systemd/system/second-boot.service
            [Unit]
            Description=Run a script only on the second boot
            Before=multi-user.target
            Wants=network.target
            After=network.target

            [Service]
            Type=oneshot
            ExecStart=/usr/local/bin/second-boot-script.sh
            RemainAfterExit=yes

            [Install]
            WantedBy=multi-user.target
            EOF
        - echo "Reloading systemd daemon and enabling second-boot service"
        - systemctl daemon-reload
        - systemctl enable second-boot.service

main.tf(片段)

resource "vsphere_virtual_machine" "hcv_nodes" {
  count            = 1
  name             = "cdavault001-0${count.index + 1}.domain.com"
  resource_pool_id = data.vsphere_compute_cluster.target_compute_cluster.resource_pool_id
  datastore_id     = data.vsphere_datastore.target_datastore.id

  num_cpus = 2
  memory   = 4096

  network_interface {
    network_id = data.vsphere_network.target_vsphere_network.id
    adapter_type = "vmxnet3"
  }

  disk {
    label            = "disk0"
    size             = 50
    eagerly_scrub    = false
    thin_provisioned = true
  }

  clone {
    template_uuid = data.vsphere_content_library_item.target_template.id

    customize {
      linux_options {
        host_name = "cdavault001-0${count.index + 1}"
        domain    = "domain"
      }

      network_interface {
        ipv4_address = "10.10.10.${31 + count.index}"
        ipv4_netmask = 24
      }

      ipv4_gateway = "10.10.10.254"

      dns_server_list = ["10.10.10.1", "8.8.8.8"]
      dns_suffix_list = ["domain.com", "domain.local"]
    }
  }
}

系统日志(/var/log/vmware-imc/toolsDeployPkg.log)

2025-01-07T12:37:36 DEBUG: Command: 'rm -rf /var/lib/dhcp/*' 
2025-01-07T12:37:36 DEBUG: Exit Code: 0 
2025-01-07T12:37:36 DEBUG: Result:  
2025-01-07T12:37:36 DEBUG: Check if command [hostnamectl] is available 
2025-01-07T12:37:36 INFO: Check if hostnamectl is available 
2025-01-07T12:37:36 DEBUG: Command: 'hostnamectl status 2>/tmp/guest.customization.stderr' 
2025-01-07T12:37:36 DEBUG: Exit Code: 1 
2025-01-07T12:37:36 DEBUG: Result:  
2025-01-07T12:37:36 DEBUG: Stderr: Failed to query system properties: Transaction for systemd-hostnamed.service/start is destructive (system-systemd\x2dfsck.slice has 'stop' job queued, but 'start' is included in transaction).
 
2025-01-07T12:37:37 INFO: Check if hostnamectl is available 
2025-01-07T12:37:37 DEBUG: Command: 'hostnamectl status 2>/tmp/guest.customization.stderr' 
2025-01-07T12:37:37 DEBUG: Exit Code: 1 
2025-01-07T12:37:37 DEBUG: Result:  
2025-01-07T12:37:37 DEBUG: Stderr: Failed to query system properties: Transaction for systemd-hostnamed.service/start is destructive (time-set.target has 'stop' job queued, but 'start' is included in transaction).
 
2025-01-07T12:37:38 INFO: Check if hostnamectl is available 
2025-01-07T12:37:38 DEBUG: Command: 'hostnamectl status 2>/tmp/guest.customization.stderr' 
2025-01-07T12:37:38 DEBUG: Exit Code: 1 
2025-01-07T12:37:38 DEBUG: Result:  
2025-01-07T12:37:38 DEBUG: Stderr: Failed to query system properties: Transaction for systemd-hostnamed.service/start is destructive (systemd-reboot.service has 'start' job queued, but 'stop' is included in transaction).
 
2025-01-07T12:37:39 INFO: Check if hostnamectl is available 
2025-01-07T12:37:39 DEBUG: Command: 'hostnamectl status 2>/tmp/guest.customization.stderr' 

=================== Perl script log end =================
[2025-01-07T12:37:39.616Z] [   error] Customization command failed with exitcode: 127, stderr: ''.
[2025-01-07T12:37:39.616Z] [   error] Customization process returned with error.
[2025-01-07T12:37:39.616Z] [   debug] Deployment result = 127.
[2025-01-07T12:37:39.616Z] [    info] Setting 'unknown' error status in vmx.
[2025-01-07T12:37:39.617Z] [    info] Transitioning from state 'INPROGRESS' to state 'ERRORED'.
[2025-01-07T12:37:39.617Z] [    info] ENTER STATE 'ERRORED'.
[2025-01-07T12:37:39.617Z] [    info] EXIT STATE 'INPROGRESS'.
[2025-01-07T12:37:39.617Z] [   debug] Setting deploy error: 'Deployment failed.The forked off process returned error code.'.
[2025-01-07T12:37:39.617Z] [   error] Deployment failed.The forked off process returned error code.
[2025-01-07T12:37:39.617Z] [    info] Launching cleanup.
[2025-01-07T12:37:39.617Z] [   debug] Command to exec : '/bin/rm'.
[2025-01-07T12:37:39.617Z] [    info] sizeof ProcessInternal is 56
[2025-01-07T12:37:39.618Z] [    info] Returning, pending output from stdout
[2025-01-07T12:37:39.618Z] [    info] Returning, pending output from stderr
[2025-01-07T12:37:39.718Z] [    info] Process exited normally after 0 seconds, returned 0
[2025-01-07T12:37:39.718Z] [    info] No more output from stdout
[2025-01-07T12:37:39.719Z] [    info] No more output from stderr
[2025-01-07T12:37:39.719Z] [    info] Customization command output:
''.
[2025-01-07T12:37:39.719Z] [    info] sSkipReboot: 'false', forceSkipReboot 'false'.
[2025-01-07T12:37:39.719Z] [   error] Deploy error: 'Deployment failed.The forked off process returned error code.'.
[2025-01-07T12:37:39.719Z] [   error] Package deploy failed in DeployPkg_DeployPackageFromFile
[2025-01-07T12:37:39.719Z] [   debug] ## Closing log

解决方案建议

方案1:调整second-boot-service的触发时机

修改systemd服务的依赖,确保在VMware自定义工具完成配置后再执行cloud-init clean:

  1. 修改服务的After参数,添加vmware-tools.service
  2. 在脚本中增加检测VMware自定义完成的标志文件逻辑,等待自定义完成后再执行清理

修改后的second-boot.service示例:

[Unit]
Description=Run a script only after VMware customization completes
Before=multi-user.target
Wants=network.target vmware-tools.service
After=network.target vmware-tools.service

[Service]
Type=oneshot
ExecStart=/usr/local/bin/second-boot-script.sh
RemainAfterExit=yes

[Install]
WantedBy=multi-user.target

脚本新增等待逻辑:

# 等待VMware自定义完成
while [ ! -f /var/lib/vmware-tools/customization_done ]; do
    echo "Waiting for VMware customization to complete..." >> /var/log/second-boot.log
    sleep 10
done

# 原有的启动次数检测与清理逻辑

方案2:在Terraform中添加等待逻辑

使用time_sleep资源,在VM克隆后等待足够时间让系统完成初始化:

resource "time_sleep" "wait_for_vm_init" {
  count = length(vsphere_virtual_machine.hcv_nodes)
  create_duration = "5m" # 根据实际情况调整等待时长

  depends_on = [vsphere_virtual_machine.hcv_nodes]
}

内容的提问来源于stack exchange,提问作者user2533368

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 04:14:49