Docker Swarm环境下Pgpool与PostgreSQL连接异常排查求助
问题排查请求
现有本地部署的Debian 12系统5节点Docker Swarm集群(3个master节点、2个worker节点),应用容器(前端+后端)运行于master节点,Pgpool集群由1个Pgpool容器和2个PostgreSQL容器组成——Pgpool容器运行在master节点,PostgreSQL容器部署在worker节点。应用整体运行正常,但偶尔后端会抛出数据库连接超时相关异常;同时Pgpool会出现write on backend 1 failed with error :"Connection reset by peer"日志,随后Pgpool容器重启。PostgreSQL容器未发生重启且日志无异常。以下为相关报错日志、Pgpool及PostgreSQL的Docker Stack配置,恳请协助排查:
后端报错日志
[08:36:53 ERR] An error occurred using the connection to database 'sampler_db' on server 'tcp://pgpool:5432'. [08:36:53 ERR] An exception occurred while iterating over the results of a query for context type 'Softgent.Sampler.Infrastructure.Data.AppDbContext'. System.InvalidOperationException: An exception has been raised that is likely due to a transient failure. ---> Npgsql.NpgsqlException (0x80004005): Exception while reading from stream ---> System.TimeoutException: Timeout during reading attempt ... System.InvalidOperationException: An exception has been raised that is likely due to a transient failure. ---> Npgsql.NpgsqlException (0x80004005): Exception while reading from stream ---> System.TimeoutException: Timeout during reading attempt ... --- End of inner exception stack trace --- at Npgsql.EntityFrameworkCore.PostgreSQL.Storage.Internal.NpgsqlExecutionStrategy.ExecuteAsync[TState,TResult](TState state, Func`4 operation, Func`4 verifySucceeded, CancellationToken cancellationToken) at Microsoft.EntityFrameworkCore.Query.Internal.SingleQueryingEnumerable`1.AsyncEnumerator.MoveNextAsync() [08:36:53 INF] Executed endpoint 'HTTP: GET /api/orders'
Pgpool报错日志
2025-02-18 08:55:47.576: psql pid 196: WARNING: write on backend 1 failed with error :"Connection reset by peer" 2025-02-18 08:55:47.576: psql pid 196: DETAIL: while trying to write data from offset: 0 wlen: 5 2025-02-18 08:55:47.578: psql pid 190: WARNING: write on backend 1 failed with error :"Connection reset by peer" 2025-02-18 08:55:47.578: psql pid 190: DETAIL: while trying to write data from offset: 0 wlen: 5 2025-02-18 08:55:47.581: psql pid 196: LOG: received degenerate backend request for node_id: 1 from pid [196] 2025-02-18 08:55:47.582: psql pid 190: LOG: received degenerate backend request for node_id: 1 from pid [190] 2025-02-18 08:55:47.582: psql pid 190: LOG: signal_user1_to_parent_with_reason(0) 2025-02-18 08:55:47.582: psql pid 196: LOG: signal_user1_to_parent_with_reason(0) 2025-02-18 08:55:47.582: main pid 1: LOG: Pgpool-II parent process received SIGUSR1 2025-02-18 08:55:47.582: psql pid 196: LOG: unable to flush data to backend 2025-02-18 08:55:47.582: psql pid 196: DETAIL: do not failover because I am the main process 2025-02-18 08:55:47.582: main pid 1: LOG: Pgpool-II parent process has received failover request 2025-02-18 08:55:47.583: main pid 1: LOG: === Starting degeneration. shutdown host postgres-1(5432) === 2025-02-18 08:55:47.583: psql pid 190: LOG: unable to flush data to backend 2025-02-18 08:55:47.583: psql pid 190: DETAIL: do not failover because I am the main process 2025-02-18 08:55:47.585: main pid 1: LOG: Do not restart children because we are switching over node id 1 host: postgres-1 port: 5432 and we are in streaming replication mode 2025-02-18 08:55:47.585: main pid 1: LOG: child pid 167 needs to restart because pool 0 uses backend 1 2025-02-18 08:55:47.585: main pid 1: LOG: child pid 180 needs to restart because pool 0 uses backend 1 2025-02-18 08:55:47.586: main pid 1: LOG: child pid 190 needs to restart because pool 0 uses backend 1 2025-02-18 08:55:47.586: main pid 1: LOG: child pid 196 needs to restart because pool 0 uses backend 1 2025-02-18 08:55:47.586: main pid 1: LOG: execute command: echo ">>> Failover - that will initialize new primary node search!" >>> Failover - that will initialize new primary node search! 2025-02-18 08:55:47.590: main pid 1: LOG: find_primary_node_repeatedly: waiting for finding a primary node 2025-02-18 08:55:47.601: main pid 1: LOG: find_primary_node: standby node is 0 2025-02-18 08:55:47.602: main pid 1: LOG: Pgpool-II parent process received SIGUSR1 2025-02-18 08:55:47.602: main pid 1: LOG: reaper handler 2025-02-18 08:55:47.602: main pid 1: LOG: reaper handler 2025-02-18 08:55:48.614: main pid 1: LOG: find_primary_node: standby node is 0 ... (the same logs: reaper handler and find_primary_node) 2025-02-18 08:56:55.337: main pid 1: LOG: exit handler called (signal: 15) 2025-02-18 08:56:55.337: main pid 1: LOG: shutting down by signal 15 2025-02-18 08:56:55.337: main pid 1: LOG: terminating all child processes
Docker PostgreSQL Stack配置
services: postgres-0: image: docker.io/bitnami/postgresql-repmgr:16-debian-12 environment: - 'POSTGRESQL_POSTGRES_PASSWORD=${POSTGRES_PASSWORD}' - 'POSTGRESQL_USERNAME=${PGPOOL_USERNAME}' - 'POSTGRESQL_PASSWORD=${PGPOOL_PASSWORD}' - 'POSTGRESQL_DATABASE=${POSTGRES_DATABASE}' - 'REPMGR_USERNAME=${REPMGR_USERNAME}' - 'REPMGR_PASSWORD=${REPMGR_PASSWORD}' - 'REPMGR_PRIMARY_HOST=${PRIMARY_NAME}' - 'REPMGR_PARTNER_NODES=${SECONDARY_NAME},${PRIMARY_NAME}' - 'REPMGR_NODE_NAME=${PRIMARY_NAME}' - 'REPMGR_NODE_NETWORK_NAME=${PRIMARY_NAME}' - 'REPMGR_NODE_ID=1' - 'POSTGRESQL_LOG_CONNECTIONS=on' - 'POSTGRESQL_LOG_DISCONNECTIONS=on' - 'BITNAMI_DEBUG=true' volumes: - /data/postgres:/bitnami/postgresql ports: - "5432" networks: traefik_public: aliases: - postgres-server-0 deploy: placement: constraints: - node.labels.type == db-master postgres-1: image: docker.io/bitnami/postgresql-repmgr:16-debian-12 environment: - 'POSTGRESQL_POSTGRES_PASSWORD=${POSTGRES_PASSWORD}' - 'POSTGRESQL_USERNAME=${PGPOOL_USERNAME}' - 'POSTGRESQL_PASSWORD=${PGPOOL_PASSWORD}' - 'POSTGRESQL_DATABASE=${POSTGRES_DATABASE}' - 'REPMGR_USERNAME=${REPMGR_USERNAME}' - 'REPMGR_PASSWORD=${REPMGR_PASSWORD}' - 'REPMGR_PRIMARY_HOST=${PRIMARY_NAME}' - 'REPMGR_PARTNER_NODES=${PRIMARY_NAME},${SECONDARY_NAME}' - 'REPMGR_NODE_NAME=${SECONDARY_NAME}' - 'REPMGR_NODE_NETWORK_NAME=${SECONDARY_NAME}' - 'REPMGR_NODE_ID=2' - 'POSTGRESQL_LOG_DISCONNECTIONS=on' - 'POSTGRESQL_LOG_CONNECTIONS=on' - 'BITNAMI_DEBUG=true' volumes: - /data/postgres:/bitnami/postgresql ports: - "5432" networks: traefik_public: aliases: - postgres-server-1 deploy: placement: constraints: - node.labels.type == db-slave pgpool: image: docker.io/bitnami/pgpool:4.5.4 labels: logging.promtail: "true" ports: - "5432:5432" configs: - source: pgpool-config target: /var/pgpool-custom.conf environment: - 'PGPOOL_USER_CONF_FILE=/var/pgpool-custom.conf' - 'PGPOOL_BACKEND_NODES=0:${PRIMARY_NAME}:5432,1:${SECONDARY_NAME}:5432' - 'PGPOOL_BACKEND_APPLICATION_NAMES=postgres-0,postgres-1' - 'PGPOOL_SR_CHECK_USER=${REPMGR_USERNAME}' - 'PGPOOL_SR_CHECK_PASSWORD=${REPMGR_PASSWORD}' - 'PGPOOL_ENABLE_LDAP=no' - 'PGPOOL_POSTGRES_USERNAME=${POSTGRES_USERNAME}' - 'PGPOOL_POSTGRES_PASSWORD=${POSTGRES_PASSWORD}' - 'PGPOOL_ADMIN_USERNAME=${PG_POOL_ADMIN_USERNAME}' - 'PGPOOL_ADMIN_PASSWORD=${PG_POOL_ADMIN_PASSWORD}' - 'PGPOOL_ENABLE_LOAD_BALANCING=yes' - 'PGPOOL_POSTGRES_CUSTOM_USERS=${PGPOOL_USERNAME}' - 'PGPOOL_POSTGRES_CUSTOM_PASSWORDS=${PGPOOL_PASSWORD}' - 'PGPOOL_AUTO_FAILBACK=yes' healthcheck: test: ["CMD", "/opt/bitnami/scripts/pgpool/healthcheck.sh"] interval: 10s timeout: 5s retries: 5 deploy: placement: constraints: - node.role == manager replicas: 1 labels: - "logging.promtail=true" networks: traefik_public: aliases: - pgpool-server networks: traefik_public: external: true configs: pgpool-config: file: ./configs/pgpool/pgpool-custom.conf
Pgpool自定义配置文件
failover_on_backend_error='on'
排查建议
- 网络层面排查:
Connection reset by peer通常是网络连接被主动断开,检查Pgpool所在master节点与PostgreSQL所在worker节点之间的网络稳定性,比如是否存在防火墙规则临时阻断、网络波动、Docker Swarm overlay网络的数据包丢失情况。可以在节点间持续ping或者用tcptrace抓包分析连接断开时的网络状态。 - 连接超时配置调整:
- 在Pgpool配置中增加
backend_connect_timeout参数(比如设置为10000毫秒),避免因后端连接超时导致的异常; - 调整应用侧的Npgsql连接超时参数,比如在EF Core配置中设置
CommandTimeout和ConnectionTimeout,适配Pgpool的故障切换窗口。
- 在Pgpool配置中增加
- Pgpool故障切换逻辑优化:
- 当前配置
failover_on_backend_error='on'会在后端出现错误时触发故障切换,但PostgreSQL本身未重启,可能是短暂的网络中断被误判。可以尝试设置failover_on_backend_error='off',同时启用backend_clustering_mode='streaming_replication'配合sr_check_period参数,让Pgpool通过流复制状态判断后端可用性,减少误触发。 - 检查
pgpool的health_check_timeout和health_check_retry_times参数,默认值可能过严,适当放宽阈值,避免因短暂的健康检查失败导致节点被剔除。
- 当前配置
- PostgreSQL连接状态检查:虽然PostgreSQL容器未重启,但可以检查其
pg_stat_activity视图,查看是否存在大量闲置连接被PostgreSQL主动回收(比如idle_in_transaction_session_timeout设置过短),导致Pgpool的连接失效。 - Docker资源限制:检查Pgpool和PostgreSQL容器所在节点的CPU、内存资源使用情况,是否存在资源耗尽导致的连接处理异常。
内容的提问来源于stack exchange,提问作者ChoosenOne
相关产品推荐
相关产品推荐

