You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决裸金属K8S集群部署RabbitMQ Operator及集群时出现的epmd nxdomain域名不存在错误

Troubleshooting RabbitMQ Cluster Deployment Failure on Bare-Metal Kubernetes

Alright, let's work through your RabbitMQ deployment issue and answer each of your questions clearly:

1. Where can I find the erl_crash.dump file in a Kubernetes environment?

The erl_crash.dump file is generated inside the RabbitMQ container's default working directory: /var/lib/rabbitmq. Since your pod is still running (even though it's unhealthy), you can check for it directly with:

kubectl exec -it definition-server-0 -n rabbitmq-system -- ls /var/lib/rabbitmq

To copy the file to your local machine for deeper analysis, run:

kubectl cp rabbitmq-system/definition-server-0:/var/lib/rabbitmq/erl_crash.dump ./erl_crash.dump

If the pod crashes completely, you can still retrieve the file from the terminated container by specifying the rabbitmq container explicitly:

kubectl cp rabbitmq-system/definition-server-0:/var/lib/rabbitmq/erl_crash.dump ./erl_crash.dump -c rabbitmq

Absolutely—your log makes this clear:

ERROR: epmd error for host definition-server-0.definition-nodes.rabbitmq-system: nxdomain (non-existing domain)

RabbitMQ depends on Erlang's epmd service to discover and communicate with cluster nodes, and it can't resolve your pod's fully qualified domain name (FQDN). A couple of key notes here:

  • In Kubernetes StatefulSets, pods get a stable FQDN in the format <pod-name>.<statefulset-headless-service>.<namespace>.svc.cluster.local. Your pod's expected full FQDN should be definition-server-0.definition-nodes.rabbitmq-system.svc.cluster.local (the missing .svc.cluster.local suffix in your error suggests a potential DNS configuration gap).
  • You added the hostname entry to your worker node's /etc/hosts, but your pod is scheduled on the master node (re-ctrl01.local, per the Scheduled event). Even if you'd added it to the correct node, manually editing node hosts files isn't a sustainable Kubernetes solution—internal hostname resolution should be handled by the cluster's DNS service.

3. Is the "connection refused" error caused by the epmd issue, or something else?

This is a direct result of the epmd failure. When epmd can't resolve the hostname, RabbitMQ can't finish its startup process, which means it never starts listening on the AMQP port (5672). The readiness probe fails simply because the port isn't open yet—fix the hostname resolution problem, and this error will go away on its own.

Let's target the root cause:

  1. Verify the StatefulSet's headless service exists:
    Run this to check if the definition-nodes headless service is properly created (it should show ClusterIP: None):
    kubectl get svc -n rabbitmq-system
    
  2. Check CoreDNS health:
    Kubernetes uses CoreDNS for internal DNS resolution. Ensure its pods are running and error-free:
    kubectl get pods -n kube-system -l k8s-app=kube-dns
    kubectl logs -n kube-system -l k8s-app=kube-dns
    
  3. Test DNS resolution inside the pod:
    Exec into your pod and test if it can resolve the expected FQDN:
    kubectl exec -it definition-server-0 -n rabbitmq-system -- nslookup definition-server-0.definition-nodes.rabbitmq-system.svc.cluster.local
    
    If this fails, your cluster's DNS configuration is broken—you'll need to troubleshoot CoreDNS or your cluster's network setup.
  4. Temporary workaround (if DNS is irreparable short-term):
    Add a hostAliases section to your StatefulSet manifest to inject the hostname entry directly into the pod's /etc/hosts:
    spec:
      template:
        spec:
          hostAliases:
          - ip: "10.244.0.xxx"  # Replace with your pod's actual IP from the describe output
            hostnames:
            - "definition-server-0.definition-nodes.rabbitmq-system"
            - "definition-server-0.definition-nodes.rabbitmq-system.svc.cluster.local"
    
    Apply the updated manifest with kubectl apply -f <your-statefulset-file.yaml>.

内容的提问来源于stack exchange,提问作者Sathish Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 20:47:31