CrunchyData Postgres实例因WAL占满存储故障,求归档配置方案
解决CrunchyData Postgres Operator WAL存储占满问题
紧急恢复步骤(先让故障实例恢复)
- 进入故障Pod,检查归档状态:
如果kubectl exec -it abc-postgres-instance1-c7ck-0 -- psql -U postgres -c "SELECT archiver_last_wal, archiver_state FROM pg_stat_archiver;"archiver_state显示failed,先手动触发一次WAL切换,尝试重新归档:kubectl exec -it abc-postgres-instance1-c7ck-0 -- psql -U postgres -c "SELECT pg_switch_wal();" - 清理已归档的本地WAL文件(确认归档成功后执行,避免数据丢失):
kubectl exec -it abc-postgres-instance1-c7ck-0 -- pg_archivecleanup /pgdata/pg16_wal/ $(kubectl exec -it abc-postgres-instance1-c7ck-0 -- psql -U postgres -t -c "SELECT pg_walfile_name(pg_current_wal_lsn());") - 删除故障Pod,让Operator自动重建:
kubectl delete pod abc-postgres-instance1-c7ck-0
长期配置:WAL大小限制与自动归档清理
1. 设置WAL保留与生成限制
修改Helm Chart的values配置文件,在postgres.parameters下添加以下参数:
postgres: parameters: wal_keep_size: "16GB" # 保留的已归档WAL总大小,超出则自动清理 max_wal_size: "8GB" # 触发检查点的最大WAL大小,控制WAL生成频率 min_wal_size: "2GB" # 检查点后保留的最小WAL大小,避免频繁清理
执行Helm升级使配置生效:
helm upgrade <your-release-name> crunchy-postgres-operator/crunchy-postgres-operator -f your-values.yaml
2. 启用WAL自动归档
配置外部存储归档WAL,避免本地存储被占满。以S3为例,在values中添加:
postgres: archive: enabled: true s3: bucket: "your-wal-archive-bucket" region: "cn-beijing" accessKey: "your-access-key-id" secretKey: "your-secret-access-key"
如果使用NFS归档,配置NFS路径(需确保归档目录与数据PVC分离):
postgres: archive: enabled: true nfs: path: "/nfs/wal-archive" server: "nfs-server-ip"
CrunchyData Operator会自动配置Postgres的archive_mode=on,并在WAL成功归档后清理本地WAL文件。
配置验证
- 检查参数是否生效:
kubectl exec -it <healthy-postgres-pod> -- psql -U postgres -c "SHOW wal_keep_size; SHOW max_wal_size; SHOW archive_mode;" - 监控本地WAL存储:
kubectl exec -it <healthy-postgres-pod> -- df -h /pgdata/pg16_wal
内容的提问来源于stack exchange,提问作者Ravindra Gupta
相关产品推荐
相关产品推荐

