如何让Ansible的until循环仅在耗尽所有重试后打印FAILED
问题描述
我有一个逻辑与Stack Overflow顶级答案几乎一致的Ansible任务,用于等待K8s控制平面Pod创建并就绪:
- name: Wait for all control-plane pods become created shell: "kubectl get po --namespace=kube-system --selector tier=control-plane --output=jsonpath='{.items[*].metadata.name}'" register: control_plane_pods_created until: item in control_plane_pods_created.stdout retries: 10 delay: 30 with_items: - etcd - kube-apiserver - kube-controller-manager - kube-scheduler - name: Wait for control-plane pods become ready shell: "kubectl wait --namespace=kube-system --for=condition=Ready pods --selector tier=control-plane --timeout=600s" register: control_plane_pods_ready - debug: var=control_plane_pods_ready.stdout_lines
执行时,第一个任务会在重试过程中反复打印FAILED - RETRYING日志,示例输出如下:
TASK [Wait for all control-plane pods become created] ****************************** FAILED - RETRYING: Wait all control-plane pods become created (10 retries left). FAILED - RETRYING: Wait all control-plane pods become created (9 retries left). FAILED - RETRYING: Wait all control-plane pods become created (8 retries left). changed: [localhost -> localhost] => (item=etcd) changed: [localhost -> localhost] => (item=kube-apiserver) changed: [localhost -> localhost] => (item=kube-controller-manager) changed: [localhost -> localhost] => (item=kube-scheduler) TASK [Wait for control-plane pods become ready] ******************************** changed: [localhost -> localhost] TASK [debug] ******************************************************************* ok: [localhost] => { "control_plane_pods_ready.stdout_lines": [ "pod/etcd-localhost.localdomain condition met", "pod/kube-apiserver-localhost.localdomain condition met", "pod/kube-controller-manager-localhost.localdomain condition met", "pod/kube-scheduler-localhost.localdomain condition met" ] }
实际场景中重试次数最多可达20次,大量冗余日志会占据输出空间,但这类重试属于预期行为。请问如何实现仅在所有重试次数耗尽、任务真正失败时才打印FAILED信息?
解决方案
方法1:通过环境变量全局控制重试日志
执行playbook时设置ANSIBLE_DISPLAY_RETRY_STATS环境变量,关闭重试过程中的失败提示:
ANSIBLE_DISPLAY_RETRY_STATS=False ansible-playbook your-playbook.yml
该变量会抑制重试阶段的FAILED - RETRYING日志输出,仅在任务最终超时失败时才打印错误信息。
方法2:优化任务逻辑,替换为原生Kubectl等待命令
原循环检查每个Pod的逻辑可以用kubectl wait的Exists条件替代,无需循环重试,自然避免冗余日志:
- name: Wait for all control-plane pods become created shell: "kubectl wait --namespace=kube-system --for=condition=Exists pods --selector tier=control-plane --timeout=300s" register: control_plane_pods_created - name: Wait for control-plane pods become ready shell: "kubectl wait --namespace=kube-system --for=condition=Ready pods --selector tier=control-plane --timeout=600s" register: control_plane_pods_ready - debug: var=control_plane_pods_ready.stdout_lines
kubectl wait --for=condition=Exists会直接等待所有匹配的Pod被创建,整个任务仅在超时失败时才输出错误,重试过程无冗余日志。
方法3:针对单个任务配置日志抑制
在目标任务中添加display_failed_hosts: false参数,抑制重试过程中的失败主机输出:
- name: Wait for all control-plane pods become created shell: "kubectl get po --namespace=kube-system --selector tier=control-plane --output=jsonpath='{.items[*].metadata.name}'" register: control_plane_pods_created until: item in control_plane_pods_created.stdout retries: 10 delay: 30 with_items: - etcd - kube-apiserver - kube-controller-manager - kube-scheduler display_failed_hosts: false
该参数会隐藏重试阶段的失败提示,仅当所有重试耗尽、任务真正失败时才显示错误信息。
内容的提问来源于stack exchange,提问作者Josh
相关产品推荐
相关产品推荐

