You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ansible连接Windows主机遇WinRM 400错误及重试失效问题

WinRM 400错误排查与重试机制问题解决

环境与问题背景

  • 部署架构:本地笔记本通过Docker运行Jenkins容器,容器内集成Ansible;Jenkins流水线对接企业VCenter中的Windows虚拟机(Windows 10/Server 2016及以上版本)
  • 流水线流程:
    1. 执行PLAY [revert vmware vcenter lab] playbook,执行成功
    2. 执行PLAY [copy files to lab] playbook时,部分虚拟机(vm1、vm2)出现winrm.exceptions.WinRMTransportError: Bad HTTP response returned from server. Code 400错误,vm3正常;手动重启copy阶段可恢复,但未解决根源问题
  • WinRM配置(group_var > all.yml):
ansible_user: bob
ansible_password: sponge
ansible_connection: winrm
ansible_port: 5985
ansible_winrm_scheme: http
ansible_winrm_server_cert_validation: ignore
ansible_winrm_kerberos_delegation: false
ansible_winrm_transport: ntlm
ansible_winrm_read_timeout_sec: 70
ansible_winrm_operation_timeout_sec: 60
  • 附加问题:任务已配置retries:5和until: results is not failed,但400错误发生时重试未触发;在虚拟机上执行ConfigureRemotingForAnsible.ps1和winRMbatch.bat脚本后,错误频率降低,但重试仍无效

需求咨询

a. 出现WinRM 400错误的原因是什么?如何解决?
b. 若无法彻底解决该错误,如何让任务自动重试直至成功?


解决方案

a. WinRM 400错误的原因与解决方法

常见原因

  • WinRM会话资源耗尽:虚拟机刚完成快照还原后,WinRM服务可能未完全初始化,或并发连接数超出默认限制,导致新请求被拒绝返回400错误
  • NTLM认证上下文失效:使用NTLM传输时,Ansible与Windows虚拟机之间的认证上下文可能因快照还原后系统状态变化而失效
  • WinRM服务配置异常:虽执行了配置脚本,但部分虚拟机的WinRM监听端口、认证策略可能仍存在细微配置差异,或服务未完全重启生效
  • 网络波动:本地Docker容器与企业VCenter虚拟机之间的网络延迟/丢包,导致WinRM请求报文不完整,触发400错误

解决步骤

  1. 调整WinRM并发连接限制
    在Windows虚拟机上执行PowerShell命令提升并发数:
    Set-Item -Path WSMan:\localhost\Shell\MaxConcurrentUsers -Value 10
    Set-Item -Path WSMan:\localhost\Shell\MaxShellsPerUser -Value 10
    Restart-Service WinRM
    
  2. 快照还原后添加延迟等待
    在PLAY [copy files to lab]开头添加任务,给WinRM服务足够初始化时间:
    - name: Wait for WinRM service to initialize after revert
      wait_for_connection:
        delay: 30
        timeout: 120
    
  3. 更换WinRM传输方式
    尝试将ansible_winrm_transport从ntlm改为basic(需先在Windows启用Basic认证),修改group_var > all.yml:
    ansible_winrm_transport: basic
    
    同时在Windows虚拟机执行:
    Set-Item WSMan:\localhost\Service\Auth\Basic -Value $true
    Restart-Service WinRM
    
  4. 快照还原后强制重启WinRM服务
    在还原快照的playbook中添加重启任务:
    - name: Restart WinRM service on Windows VMs
      win_service:
        name: WinRM
        state: restarted
      delegate_to: "{{ item }}"
      loop: "{{ groups['windows_lab'] }}"
    

b. 确保重试机制生效的配置方法

原配置重试未触发,是因为winrm.exceptions.WinRMTransportError属于连接级别的致命错误,默认不会触发任务级重试,可通过以下方式调整:

  1. 使用block/rescue块捕获错误并重试
    在copy任务外层包裹结构,手动捕获WinRM错误并触发重试:
    - name: Copy files to Windows VMs with retry
      block:
        - name: Copy target files
          win_copy:
            src: /path/to/local/files/
            dest: C:\target\path\
          register: results
      rescue:
        - name: Retry copy task on WinRM error
          win_copy:
            src: /path/to/local/files/
            dest: C:\target\path\
          register: retry_results
          until: retry_results is not failed
          retries: 5
          delay: 10
    
  2. 配置Ansible全局错误重试策略
    在ansible.cfg中添加配置,让Ansible自动重试指定WinRM错误:
    [defaults]
    retry_files_enabled = false
    [connection]
    retries = 3
    timeout = 60
    [winrm]
    retry_winrm_errors = 400,500,503
    
    该配置会让Ansible对WinRM返回的400、500、503错误自动重试3次
  3. 结合wait_for_connection验证连接后重试
    在重试前先确认WinRM连接可用,避免无效重试:
    - name: Copy files with connection check retry
      win_copy:
        src: /path/to/local/files/
        dest: C:\target\path\
      register: copy_result
      until: copy_result is not failed
      retries: 5
      delay: 10
      failed_when: false
      notify: Wait for WinRM connection
    handlers:
      - name: Wait for WinRM connection
        wait_for_connection:
          timeout: 60
    

内容的提问来源于stack exchange,提问作者Hiddai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 23:20:40