K8s中Pod无法连接数据库,出现名称解析错误
K8s部署服务连接PostgreSQL失败:域名解析错误排查
问题现象
在单机器上K8s部署服务与PostgreSQL数据库正常,但另一台机器部署失败。数据库Pod日志正常,但服务连接数据库时报错:
... raise exception File "/usr/local/lib/python3.8/site-packages/sqlalchemy/pool/base.py", line 656, in __connect connection = pool._invoke_creator(self) File "/usr/local/lib/python3.8/site-packages/sqlalchemy/engine/strategies.py", line 114, in connect return dialect.connect(*cargs, **cparams) File "/usr/local/lib/python3.8/site-packages/sqlalchemy/engine/default.py", line 490, in connect return self.dbapi.connect(*cargs, **cparams) File "/usr/local/lib/python3.8/site-packages/psycopg2/__init__.py", line 122, in connect conn = _connect(dsn, connection_factory=connection_factory, **kwasync) sqlalchemy.exc.OperationalError: (psycopg2.OperationalError) could not translate host name "sbdb-0.sbdb" to address: Temporary failure in name resolution
相关资源信息
服务Pod详情
lev@gpusrv01:~$ sudo kubectl describe pod sb-6d5677f566-4qmmf Name: sb-6d5677f566-4qmmf Namespace: default Priority: 0 Node: gpusrv01/10.200.3.64 Start Time: Sat, 04 Mar 2023 10:58:33 +0100 Labels: app=servicebroker-app db=sbdb name=sb-pod pod-template-hash=6d5677f566 Annotations: sidecar.istio.io/rewriteAppHTTPProbers: true Status: Running IP: 10.42.0.193 IPs: IP: 10.42.0.193 Controlled By: ReplicaSet/sb-6d5677f566 Containers: sb: Container ID: containerd://d713eb355a0dafd9dcf9b7da6bdd5b6b1cca55147846af9f68a94e595614df4b Image: sb Image ID: sha256:3dc49d5c108e81948af187d2ad00f85d6b1831b8061ca79f1f3ca44a29955054 Ports: 5000/TCP, 5000/TCP Host Ports: 0/TCP, 0/TCP Command: flask run --host=0.0.0.0 --port=5000 State: Running Started: Sat, 04 Mar 2023 10:58:33 +0100 Ready: False Restart Count: 0 Readiness: http-get http://:5000/heliports delay=0s timeout=1s period=10s #success=1 #failure=3 Environment Variables from: sb-env-list ConfigMap Optional: false Environment: FLASK_APP: app.py DB_USER: <set to the key 'DB_USER' in secret 'servicebroker-secrets'> Optional: false DB_PASSWORD: <set to the key 'DB_PASSWORD' in secret 'servicebroker-secrets'> Optional: false Mounts: <none> Conditions: Type Status Initialized True Ready False ContainersReady False PodScheduled True Volumes: <none> QoS Class: BestEffort Node-Selectors: <none> Tolerations: node.kubernetes.io/not-ready:NoExecute op=Exists for 300s node.kubernetes.io/unreachable:NoExecute op=Exists for 300s Events: Type Reason Age From Message ---- ------ ---- ---- ------- Warning Unhealthy 3m19s (x18830 over 46h) kubelet Readiness probe failed: Get "http://10.42.0.193:5000/heliports": context deadline exceeded (Client.Timeout exceeded while awaiting headers)
服务Service详情
lev@gpusrv01:~$ sudo kubectl describe svc sb Name: sb Namespace: default Labels: app=servicebroker-app name=sb Annotations: <none> Selector: app=servicebroker-app,db=sbdb,name=sb-pod Type: NodePort IP Family Policy: SingleStack IP Families: IPv4 IP: 10.43.176.221 IPs: 10.43.176.221 Port: http 5000/TCP TargetPort: 5000/TCP NodePort: http 30050/TCP Endpoints: Session Affinity: None External Traffic Policy: Cluster Events: <none>
数据库连接配置(ConfigMap)
apiVersion: v1 kind: ConfigMap metadata: name: sb-env-list data: DEMO_MODE: 'False' DB_HOST: sbdb-0.sbdb DB_NAME: servicebroker DB_PORT: "5432" LOG_LEVEL: "info" KAFKA_ADDRESS: kk-cluster-kafka-bootstrap.kk.svc.cluster.local:9092 MRM_ADDRESS: "http://mrm:5001"
排查结论
核心问题是服务Pod无法解析PostgreSQL域名sbdb-0.sbdb,大概率是DNS或集群网络配置问题,可按以下步骤验证修复:
- 确认PostgreSQL的Headless Service
sbdb存在,且关联的Podsbdb-0状态正常,sbdb-0.sbdb是StatefulSet Pod的标准域名格式(Pod名.Headless服务名) - 进入服务Pod执行
nslookup sbdb-0.sbdb,直接测试域名解析能力,若失败则说明Pod DNS配置异常 - 检查
kube-system命名空间下CoreDNS/kube-dns组件的运行状态,确认无崩溃或重启 - 排查故障节点的防火墙/iptables规则,确保未阻断53端口(DNS默认端口)的UDP/TCP请求
- 确认服务与PostgreSQL在同一命名空间,若不在,域名需补充命名空间后缀:
sbdb-0.sbdb.<命名空间>.svc.cluster.local
服务Pod的就绪探针超时是连锁问题,根源解决后探针状态会自动恢复。
内容的提问来源于stack exchange,提问作者Leviathan
相关产品推荐
相关产品推荐

