Nginx Ingress跨多Namespace查找Service引发间歇性503错误求助
Alright, let's break down what's happening here and get this sorted out step by step.
What's Causing the Problem?
When you configure Nginx to forward requests to just frontend (a short service name), Kubernetes' DNS system might be resolving that name across all namespaces if:
- Your Nginx pod's DNS search domains include broader cluster scopes (like
svc.cluster.local), which lets it find anyfrontendservice in the cluster. - Test and staging namespaces actually have a
frontendservice defined (even if there are no pods running behind it).
Nginx picks up all these resolved service addresses, adds them to its upstream pool, and rounds-robin requests across them. Since test/staging's frontend has no active pods (empty endpoints), those requests fail with 503 errors.
Step-by-Step Solutions
1. Use Fully Qualified Domain Names (FQDN) for Upstream Services
The quickest fix is to explicitly tell Nginx to only target the frontend service in the dev namespace by using its full DNS name.
In your Nginx configuration, replace lines like:
server frontend:80;
with:
server frontend.dev.svc.cluster.local:80;
This ensures DNS only resolves the dev namespace's frontend service, so Nginx won't look elsewhere.
2. Restrict Nginx Pod's DNS Search Domains
If you prefer keeping short service names, lock down the Nginx pod's DNS settings to prioritize the dev namespace first.
Update your Nginx deployment YAML to include a custom dnsConfig:
apiVersion: apps/v1 kind: Deployment metadata: name: nginx-deployment namespace: dev spec: template: spec: dnsPolicy: "None" dnsConfig: searches: - dev.svc.cluster.local - svc.cluster.local - cluster.local nameservers: - 10.96.0.10 # Default Kubernetes DNS service IP (adjust if your cluster uses a different one)
This forces the pod to only look for services in dev first, avoiding cross-namespace matches.
3. Clean Up Unnecessary Services in Test/Staging
If those frontend services in test and staging aren't serving any purpose, delete them entirely to eliminate the rogue endpoints:
kubectl delete service frontend -n test kubectl delete service frontend -n staging
4. Verify the Fix
To confirm everything works as expected, exec into your dev Nginx pod and check DNS resolution:
kubectl exec -it <your-nginx-pod-name> -n dev -- nslookup frontend kubectl exec -it <your-nginx-pod-name> -n dev -- nslookup frontend.dev.svc.cluster.local
Both commands should only return the IP address of the dev namespace's frontend service. No extra IPs from test/staging means no more random 503s.
内容的提问来源于stack exchange,提问作者Narayan Prusty

