You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PostgreSQL-HA随机出现上游节点连接失败及Pgpool存活探针异常

PostgreSQL HA(Bitnami)在Istio集群中连接波动导致Pgpool存活探针失败问题

环境:注入了Istio Sidecar的Kubernetes集群。

我使用bitnami/postgresql-ha作为Airflow的数据库,3个Pod组成的PostgreSQL StatefulSet(镜像:bitnami/postgresql-repmgr:15.3.0-debian-11-r8)会随机出现以下日志,无固定规律,有时一天出现10+次,有时仅1次:

[2023-08-18 02:41:42] [WARNING] unable to ping "user=repmgr password=admin host=airflow-postgresql-1.airflow-postgresql-headless.workflow.svc.cluster.local dbname=repmgr port=5432 connect_timeout=5"
[2023-08-18 02:41:42] [DETAIL] PQping() returned "PQPING_NO_RESPONSE"
[2023-08-18 02:41:42] [WARNING] unable to connect to upstream node "airflow-postgresql-1" (ID: 1001)
[2023-08-18 02:41:42] [NOTICE] node "airflow-postgresql-1" (ID: 1001) has recovered, reconnecting
[2023-08-18 02:41:42] [NOTICE] reconnected to upstream node after 0 seconds

注意:总能在0秒内重新连接。

该问题会导致Pgpool的存活探针失败,触发如下事件信息,进而导致Airflow任务失败:

Liveness probe failed: Checking pgpool health... 
psql: error: connection to server on socket "/opt/bitnami/pgpool/tmp/.s.PGSQL.5432" 
failed: ERROR: unable to read message kind DETAIL: kind does not match between main(0) slot[0] (52)

已尝试以下方案,但均无效:

  • 延长Pgpool存活探针的periodSeconds和timeoutSeconds;
  • 将Pgpool副本数从2改为1;
  • 在PostgreSQL中设置pgHbaTrustAll为true;
  • 更换PostgreSQL和Pgpool镜像版本(试过Pgpool 4.3、4.4,repmgr 14、15);
  • 在另一个K8s集群部署相同架构,问题仍存在;
  • 关闭Pgpool负载均衡;
  • 将最大连接数增加至10000;

已确认:所有相关Pod的CPU、内存资源充足。

内容的提问来源于stack exchange,提问作者Jasmine H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 07:36:09