Kubernetes集群Ingress访问Pihole服务返回504错误排查求助
树莓派K8s集群Ingress转发Pihole返回504错误问题及解决方案
问题背景
我在家使用多台树莓派搭建集群学习Kubernetes,已成功在集群内部署运行Pihole,当前遇到Ingress转发异常问题:访问对应域名返回504错误,从Pod内部curl集群IP可得到预期响应。相关配置和排查信息如下:
1. Ingress配置文件内容
## pihole.ingress.yml apiVersion: networking.k8s.io/v1 kind: Ingress metadata: namespace: pihole name: pihole-ingress annotations: # use the shared ingress-nginx kubernetes.io/ingress.class: "nginx" nginx.ingress.kubernetes.io/rewrite-target: / spec: rules: - host: pihole.192.168.1.230.nip.io http: paths: - path: / pathType: ImplementationSpecific backend: service: name: pihole-web port: number: 80
2. Ingress资源详情
kubectl describe ingress -n pihole 输出:
Name: pihole-ingress Namespace: pihole Address: 192.168.1.230 Default backend: default-http-backend:80 (<error: endpoints "default-http-backend" not found>) Rules: Host Path Backends ---- ---- -------- pihole.192.168.1.230.nip.io / pihole-web:80 (10.42.2.7:80) Annotations: kubernetes.io/ingress.class: nginx nginx.ingress.kubernetes.io/rewrite-target: / Events: Type Reason Age From Message ---- ------ ---- ---- ------- Normal Sync 18s (x12 over 11h) nginx-ingress-controller Scheduled for sync
3. Ingress Nginx服务信息
kubectl get svc -n ingress-nginx 输出:
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE ingress-nginx-controller-admission ClusterIP 10.43.240.186 <none> 443/TCP 22h ingress-nginx-controller LoadBalancer 10.43.64.54 192.168.1.230 80:31093/TCP,443:30179/TCP 22h
4. Pihole命名空间资源信息
kubectl get po -n pihole -o wide 与 kubectl get svc -n pihole 输出:
kubectl get po -n pihole -o wide NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES pihole-7d4dc6b8d8-vclxz 1/1 Running 0 9h 10.42.2.8 node02.iad <none> <none> kubectl get svc -n pihole NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE pihole-web ClusterIP 10.43.102.198 <none> 80/TCP,443/TCP 9h pihole-dhcp NodePort 10.43.191.110 <none> 67:32021/UDP 9h pihole-dns-udp NodePort 10.43.214.15 <none> 53:31153/UDP 9h pihole-dns-tcp NodePort 10.43.168.6 <none> 53:32754/TCP 9h
注:问题发布后Pihole Pod IP从10.42.2.7更新为10.42.2.8
5. Ingress控制器错误日志
2021/09/03 17:52:35 [error] 1938#1938: *3132346 upstream timed out (110: Operation timed out) while connecting to upstream, client: 10.42.1.1, server: pihole.192.168.1.230.nip.io, request: "GET / HTTP/1.1", upstream: "http://10.42.2.8:80/", host: "pihole.192.168.1.230.nip.io", referrer: "http://pihole.192.168.1.230.nip.io/"
问题根因与修复方案
从错误日志可以直接判断:Ingress Nginx Pod无法连接到上游的Pihole Pod,本质是集群跨节点的Pod网络连通性故障,和Ingress配置本身无关,按以下步骤排查修复即可:
- 验证跨节点Pod连通性:登录Ingress Nginx Pod所在的节点,直接执行
curl 10.42.2.8:80,如果超时即可确认是节点间Pod网段不通。 - 放行CNI插件端口:树莓派默认的防火墙规则会拦截CNI插件的跨节点通信流量,如果你使用Flannel作为CNI插件,需要放行UDP 8472端口;如果使用Calico,需要放行TCP 4789、5473端口。
- 关闭RP过滤:树莓派系统默认开启的反向路径过滤会丢弃非对称路由的Pod流量,执行以下命令临时关闭验证:
sysctl -w net.ipv4.conf.all.rp_filter=0 sysctl -w net.ipv4.conf.default.rp_filter=0
验证访问正常后,将上述两行配置写入/etc/sysctl.conf文件实现永久生效。
4. 可选优化:当前Ingress配置中nginx.ingress.kubernetes.io/rewrite-target: /对于根路径转发属于冗余配置,可以直接删除。
内容的提问来源于stack exchange,提问作者Muji Sayed
相关产品推荐
相关产品推荐

