Kubernetes集群Ingress-Nginx就绪探针失败问题求助
全新部署的Kubernetes集群中,通过Manifest清单和Helm Chart两种方式部署Ingress-Nginx,均出现Pod无法就绪的情况。
执行kubectl get po查看Pod状态,结果如下:
kubectl get po nginx-ingress-dx6bg 0/1 Running 3 (26s ago) 3m44s 10.244.4.118 node-2 <none> <none> nginx-ingress-gqkhz 0/1 Running 3 (29s ago) 3m47s 10.244.3.16 node-1 <none> <none> nginx-ingress-dx6bg 0/1 Error 3 (86s ago) 4m44s 10.244.4.118 node-2 <none> <none> nginx-ingress-gqkhz 0/1 Error 3 (89s ago) 4m47s 10.244.3.16 node-1 <none> <none> nginx-ingress-dx6bg 0/1 CrashLoopBackOff 3 (12s ago) 4m56s 10.244.4.118 node-2 <none> <none> nginx-ingress-gqkhz 0/1 CrashLoopBackOff 3 (13s ago) 4m59s 10.244.3.16 node-1 <none> <none> nginx-ingress-gqkhz 0/1 Running 4 (44s ago) 5m30s 10.244.3.16 node-1 <none> <none> nginx-ingress-dx6bg 0/1 Running 4 (51s ago) 5m35s 10.244.4.118 node-2 <none> <none> nginx-ingress-b9fcfbb59-hwjc8 0/1 Running 6 (2m49s ago) 12m 10.244.4.116 node-2 <none> <none>
通过kubectl describe po -n nginx-ingress nginx-ingress-b9fcfbb59-hwjc8查看Pod详情,发现就绪探针持续失败,事件提示:
Warning Unhealthy 3m56s (x250 over 8m34s) kubelet Readiness probe failed: Get "http://10.244.4.116:8081/nginx-ready": dial tcp 10.244.4.116:8081: connect: connection refused
Pod完整详情输出:
kd po -n nginx-ingress nginx-ingress-b9fcfbb59-hwjc8 Name: nginx-ingress-b9fcfbb59-hwjc8 Namespace: nginx-ingress Priority: 0 Service Account: nginx-ingress Node: node-2/192.168.17.15 Start Time: Thu, 08 Feb 2024 17:09:37 +0100 Labels: app=nginx-ingress app.kubernetes.io/name=nginx-ingress app.kubernetes.io/version=3.4.2 app.nginx.org/version=1.25.3 pod-template-hash=b9fcfbb59 Annotations: <none> Status: Running SeccompProfile: RuntimeDefault IP: 10.244.4.116 IPs: IP: 10.244.4.116 Controlled By: ReplicaSet/nginx-ingress-b9fcfbb59 Containers: nginx-ingress: Container ID: containerd://57299408237d9d8b1b7be67ac12d6999640ff2249305c8d289a78a58fe6b38c9 Image: nginx/nginx-ingress:3.4.2 Image ID: docker.io/nginx/nginx-ingress@sha256:4b97f1d3466c804d51abbdeb84f2c7c3ea00d6a937a320d62a4cf9d6b447d6ad Ports: 80/TCP, 443/TCP, 8081/TCP, 9113/TCP Host Ports: 0/TCP, 0/TCP, 0/TCP, 0/TCP Args: -nginx-configmaps=$(POD_NAMESPACE)/nginx-config State: Running Started: Thu, 08 Feb 2024 17:17:51 +0100 Last State: Terminated Reason: Error Exit Code: 255 Started: Thu, 08 Feb 2024 17:15:30 +0100 Finished: Thu, 08 Feb 2024 17:16:30 +0100 Ready: False Restart Count: 5 Requests: cpu: 100m memory: 128Mi Readiness: http-get http://:readiness-port/nginx-ready delay=0s timeout=1s period=1s #success=1 #failure=3 Environment: POD_NAMESPACE: nginx-ingress (v1:metadata.namespace) POD_NAME: nginx-ingress-b9fcfbb59-hwjc8 (v1:metadata.name) Mounts: /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-vlfd8 (ro) Conditions: Type Status Initialized True Ready False ContainersReady False PodScheduled True Volumes: kube-api-access-vlfd8: Type: Projected (a volume that contains injected data from multiple sources) TokenExpirationSeconds: 3607 ConfigMapName: kube-root-ca.crt ConfigMapOptional: <nil> DownwardAPI: true QoS Class: Burstable Node-Selectors: <none> Tolerations: node.kubernetes.io/not-ready:NoExecute op=Exists for 300s node.kubernetes.io/unreachable:NoExecute op=Exists for 300s Events: Type Reason Age From Message ---- ------ ---- ---- ------- Normal Scheduled 8m57s default-scheduler Successfully assigned nginx-ingress/nginx-ingress-b9fcfbb59-hwjc8 to node-2 Normal Pulling 8m57s kubelet Pulling image "nginx/nginx-ingress:3.4.2" Normal Pulled 8m35s kubelet Successfully pulled image "nginx/nginx-ingress:3.4.2" in 21.588s (21.589s including waiting) Normal Created 8m35s kubelet Created container nginx-ingress Normal Started 8m35s kubelet Started container nginx-ingress Warning Unhealthy 3m56s (x250 over 8m34s) kubelet Readiness probe failed: Get "http://10.244.4.116:8081/nginx-ready": dial tcp 10.244.4.116:8081: connect: connection refused
已尝试通过Helm设置nginxReloadTimeout=20000,但无效果,Helm安装命令:
helm install nginx-ingress-controller nginx-stable/nginx-ingress --set rbac.create=true --set controller."nodeSelector\.kubernetes\.io/hostname"=node-2 --set nginxReloadTimeout=20000
该Ingress-Nginx版本在其他集群可正常运行,寻求无需重置集群的解决建议。
解决建议
检查Pod内NGINX启动日志:进入Pod内部查看NGINX错误日志,确认启动失败原因,比如配置文件错误、权限问题:
kubectl logs -n nginx-ingress nginx-ingress-b9fcfbb59-hwjc8 kubectl exec -n nginx-ingress nginx-ingress-b9fcfbb59-hwjc8 -- cat /var/log/nginx/error.log验证指定的ConfigMap是否存在:Pod启动参数中指定了
-nginx-configmaps=$(POD_NAMESPACE)/nginx-config,检查该ConfigMap是否存在,不存在则创建空ConfigMap或移除该启动参数:kubectl get configmap -n nginx-ingress nginx-config # 若不存在,创建空ConfigMap kubectl create configmap nginx-config -n nginx-ingress调整就绪探针参数:当前探针
initialDelaySeconds=0,容器启动后立即探测,可能NGINX未完全就绪。通过Helm调整参数:helm upgrade nginx-ingress-controller nginx-stable/nginx-ingress \ --set controller.readinessProbe.initialDelaySeconds=10 \ --set controller.readinessProbe.timeoutSeconds=5 \ --set controller.readinessProbe.periodSeconds=5Manifest部署则直接修改Pod模板的
readinessProbe字段。排查Pod网络连通性:在Pod内尝试访问本地就绪端口,确认是否是NGINX未监听端口,还是集群网络问题:
kubectl exec -n nginx-ingress nginx-ingress-b9fcfbb59-hwjc8 -- curl http://localhost:8081/nginx-ready若本地无法访问,说明NGINX启动异常;若本地可访问,检查CNI插件(如Calico/Flannel)状态或节点防火墙规则。
检查ServiceAccount权限:确认Ingress-Nginx使用的ServiceAccount有访问Kubernetes API的权限:
kubectl auth can-i get configmaps -n nginx-ingress --as=system:serviceaccount:nginx-ingress:nginx-ingress若返回
no,补充RBAC规则或确认Helm安装时rbac.create=true已生效。更换Ingress-Nginx版本:新集群可能存在版本兼容性问题,尝试稳定旧版本:
helm install nginx-ingress-controller nginx-stable/nginx-ingress \ --version 3.3.0 \ --set rbac.create=true \ --set controller."nodeSelector\.kubernetes\.io/hostname"=node-2
内容的提问来源于stack exchange,提问作者Roberto D. Maggi

