OpenShift 3.11中MongoDB 7.0.2配置探针后副本集异常
OpenShift中单节点MongoDB副本集探针配置问题排查与修复
问题背景
- 因备份要求,需在OpenShift中部署单节点MongoDB副本集,Dockerfile使用
entrypoint.sh而非CMD - 未配置探针时部署正常,添加存活/就绪探针后出现两个核心问题:
- 存活探针导致Pod删除后数据库无法恢复
- 就绪探针始终失效,容器永久处于未就绪状态
- 副本集初始化命令:
rs.initiate( { _id: 'rs0', members: [ { _id: 0, host: '(openshift-service-name):27017'}, ] });
- 当前探针配置:
livenessProbe: exec: command: [ "mongosh", "--host", "localhost", "--port", "27017", "-u", "admin", "-p", "'admin'", "--authenticationDatabase", "'admin'", "--eval", "'db.getSiblingDB(\"admin\").runCommand({ replSetGetStatus: 1 }).ok ? 0 : 2'" ] initialDelaySeconds: 70 periodSeconds: 60 failureThreshold: 3 timeoutSeconds: 20 readinessProbe: exec: command: [ "mongosh", "--host", "localhost", "--port", "27017", "-u", "admin", "-p", "'admin'", "--authenticationDatabase", "'admin'", "--eval", "'db.runCommand({ ping: 1 }).ok ? 0 : 2'" ] initialDelaySeconds: 70 periodSeconds: 60 failureThreshold: 3 timeoutSeconds: 20
- 关键错误日志:
{"t":{"$date":"2024-01-18T06:15:34.306+00:00"},"s":"I", "c":"CONTROL", "id":20711, "ctx":"LogicalSessionCacheReap","msg":"Failed to reap transaction table","attr":{"error":"NotYetInitialized: Replication has not yet been configured"}} {"t":{"$date":"2024-01-18T06:15:34.306+00:00"},"s":"I", "c":"SHARDING", "id":7012500, "ctx":"QueryAnalysisConfigurationsRefresher","msg":"Failed to refresh query analysis configurations, will try again at the next interval","attr":{"error":"PrimarySteppedDown: No primary exists currently"}} {"t":{"$date":"2024-01-18T06:15:36.800+00:00"},"s":"I", "c":"REPL", "id":21394, "ctx":"ReplCoord-0","msg":"This node is not a member of the config"} {"t":{"$date":"2024-01-18T06:15:36.800+00:00"},"s":"I", "c":"REPL", "id":21358, "ctx":"ReplCoord-0","msg":"Replica set state transition","attr":{"newState":"REMOVED","oldState":"STARTUP"}}
问题分析
1. 存活探针导致数据库恢复失败
日志中This node is not a member of the config和newState: REMOVED表明:Pod重启后MongoDB节点无法重新加入副本集。原因是:
- 存活探针依赖
replSetGetStatus命令,副本集未完全初始化时探针会判定容器不健康并重启Pod,破坏副本集配置 - 单节点副本集重启后,若节点被标记为
REMOVED,当前entrypoint.sh未处理自动恢复逻辑
2. 就绪探针失效
核心问题是探针命令的认证参数错误:
- 密码和认证数据库被额外添加单引号(
-p "'admin'"),mongosh会将单引号视为密码的一部分,导致认证失败 - 初始延迟70秒可能不足,单节点副本集初始化在资源受限环境中需要更长时间
修复方案
第一步:修正探针命令的认证参数
移除多余单引号,调整探针时间参数:
livenessProbe: exec: command: - mongosh - --host - localhost - --port - "27017" - -u - admin - -p - admin - --authenticationDatabase - admin - --eval - "db.getSiblingDB('admin').runCommand({ replSetGetStatus: 1 }).ok ? 0 : 2" initialDelaySeconds: 120 periodSeconds: 30 failureThreshold: 5 timeoutSeconds: 10 readinessProbe: exec: command: - mongosh - --host - localhost - --port - "27017" - -u - admin - -p - admin - --authenticationDatabase - admin - --eval - "db.runCommand({ ping: 1 }).ok ? 0 : 2" initialDelaySeconds: 90 periodSeconds: 15 failureThreshold: 3 timeoutSeconds: 10
第二步:修改entrypoint.sh添加副本集恢复逻辑
在entrypoint.sh中添加检测逻辑,自动处理重启后的副本集恢复:
#!/bin/bash set -e # 启动MongoDB后台进程 mongod --replSet rs0 --bind_ip_all --setParameter enableMajorityReadConcern=false & MONGOD_PID=$! # 等待MongoDB启动完成 sleep 15 # 检测并修复副本集状态 mongosh --host localhost:27017 -u admin -p admin --authenticationDatabase admin --eval " try { const status = rs.status(); const currentNode = '$(openshift-service-name):27017'; // 若当前节点被标记为REMOVED,强制重新配置副本集 if (status.members.some(m => m.name === currentNode && m.stateStr === 'REMOVED')) { rs.reconfig({_id: 'rs0', members: [{_id: 0, host: currentNode}]}, {force: true}); print('Reinitialized replica set after node removal'); } } catch (e) { // 若副本集未初始化,执行初始化 if (e.message.includes('NotYetInitialized')) { rs.initiate({_id: 'rs0', members: [{_id: 0, host: '$(openshift-service-name):27017'}]}); print('Initialized new replica set'); } else { throw e; } } " # 持续监听MongoDB进程 wait $MONGOD_PID
注意:将
$(openshift-service-name)替换为实际的OpenShift服务名称
第三步:关闭多数读关注
启动MongoDB时添加--setParameter enableMajorityReadConcern=false,解决日志中ReadConcernMajorityNotAvailableYet错误,该参数已包含在上述entrypoint的启动命令中。
内容的提问来源于stack exchange,提问作者Diegunio
相关产品推荐
相关产品推荐

