如何解决裸金属K8S集群部署RabbitMQ Operator及集群时出现的epmd nxdomain域名不存在错误
Alright, let's work through your RabbitMQ deployment issue and answer each of your questions clearly:
1. Where can I find the erl_crash.dump file in a Kubernetes environment?
The erl_crash.dump file is generated inside the RabbitMQ container's default working directory: /var/lib/rabbitmq. Since your pod is still running (even though it's unhealthy), you can check for it directly with:
kubectl exec -it definition-server-0 -n rabbitmq-system -- ls /var/lib/rabbitmq
To copy the file to your local machine for deeper analysis, run:
kubectl cp rabbitmq-system/definition-server-0:/var/lib/rabbitmq/erl_crash.dump ./erl_crash.dump
If the pod crashes completely, you can still retrieve the file from the terminated container by specifying the rabbitmq container explicitly:
kubectl cp rabbitmq-system/definition-server-0:/var/lib/rabbitmq/erl_crash.dump ./erl_crash.dump -c rabbitmq
2. Is this definitely a hostname-related issue?
Absolutely—your log makes this clear:
ERROR: epmd error for host definition-server-0.definition-nodes.rabbitmq-system: nxdomain (non-existing domain)
RabbitMQ depends on Erlang's epmd service to discover and communicate with cluster nodes, and it can't resolve your pod's fully qualified domain name (FQDN). A couple of key notes here:
- In Kubernetes StatefulSets, pods get a stable FQDN in the format
<pod-name>.<statefulset-headless-service>.<namespace>.svc.cluster.local. Your pod's expected full FQDN should bedefinition-server-0.definition-nodes.rabbitmq-system.svc.cluster.local(the missing.svc.cluster.localsuffix in your error suggests a potential DNS configuration gap). - You added the hostname entry to your worker node's
/etc/hosts, but your pod is scheduled on the master node (re-ctrl01.local, per theScheduledevent). Even if you'd added it to the correct node, manually editing node hosts files isn't a sustainable Kubernetes solution—internal hostname resolution should be handled by the cluster's DNS service.
3. Is the "connection refused" error caused by the epmd issue, or something else?
This is a direct result of the epmd failure. When epmd can't resolve the hostname, RabbitMQ can't finish its startup process, which means it never starts listening on the AMQP port (5672). The readiness probe fails simply because the port isn't open yet—fix the hostname resolution problem, and this error will go away on its own.
Recommended Fixes
Let's target the root cause:
- Verify the StatefulSet's headless service exists:
Run this to check if thedefinition-nodesheadless service is properly created (it should showClusterIP: None):kubectl get svc -n rabbitmq-system - Check CoreDNS health:
Kubernetes uses CoreDNS for internal DNS resolution. Ensure its pods are running and error-free:kubectl get pods -n kube-system -l k8s-app=kube-dns kubectl logs -n kube-system -l k8s-app=kube-dns - Test DNS resolution inside the pod:
Exec into your pod and test if it can resolve the expected FQDN:
If this fails, your cluster's DNS configuration is broken—you'll need to troubleshoot CoreDNS or your cluster's network setup.kubectl exec -it definition-server-0 -n rabbitmq-system -- nslookup definition-server-0.definition-nodes.rabbitmq-system.svc.cluster.local - Temporary workaround (if DNS is irreparable short-term):
Add ahostAliasessection to your StatefulSet manifest to inject the hostname entry directly into the pod's/etc/hosts:
Apply the updated manifest withspec: template: spec: hostAliases: - ip: "10.244.0.xxx" # Replace with your pod's actual IP from the describe output hostnames: - "definition-server-0.definition-nodes.rabbitmq-system" - "definition-server-0.definition-nodes.rabbitmq-system.svc.cluster.local"kubectl apply -f <your-statefulset-file.yaml>.
内容的提问来源于stack exchange,提问作者Sathish Kumar

