MongoDB 3.2.17分片集群停止Balancer时旧ping日期告警是否正常?
问题背景
集群架构:3个分片(Shard),每个分片为3节点副本集(Replica Set),配置服务器(Config Server)也采用副本集架构。
我们在备份脚本中执行停止Balancer操作时,出现如下提示信息,无法确认是否为正常现象:
Ping显示旧日期(识别新配置阶段),可能存在正在进行的迁移或主机已宕机
Stopping Balancer .... Waiting for active hosts... Waiting for active host node-mongo3-3:27017 to recognize new settings... (ping : Sat Sep 04 2021 15:53:14 GMT+0000 (GMT)) Waited for active ping to change for host node-mongo3-3:27017, a migration may be in progress or the host may be down. Waiting for active host node-mongo2-3:27017 to recognize new settings... (ping : Mon Oct 18 2021 09:05:00 GMT+0000 (GMT)) Waiting for active host node-mongo1-3:27017 to recognize new settings... (ping : Tue Aug 31 2021 15:28:27 GMT+0000 (GMT)) Waited for active ping to change for host node-mongo1-3:27017, a migration may be in progress or the host may be down. Waiting for the balancer lock... Waiting again for active hosts after balancer is off... Waiting for active host node-mongo3-3:27017 to recognize new settings... (ping : Sat Sep 04 2021 15:53:14 GMT+0000 (GMT)) Waited for active ping to change for host node-mongo3-3:27017, a migration may be in progress or the host may be down. Waiting for active host node-mongo1-3:27017 to recognize new settings... (ping : Tue Aug 31 2021 15:28:27 GMT+0000 (GMT)) Waited for active ping to change for host node-mongo1-3:27017, a migration may be in progress or the host may be down. Warning : host node-mongo3-3:27017 seems to have been offline since Sat Sep 04 2021 15:53:14 GMT+0000 (GMT) Warning : host node-mongo1-3:27017 seems to have been offline since Tue Aug 31 2021 15:28:27 GMT+0000 (GMT) WriteResult({ "nMatched" : 1, "nUpserted" : 0, "nModified" : 1 }) Checking if the balancer is stopped Balancer is not running now. --------- 在mongos配置查询中也可见旧ping日期: db.mongos.find() { "_id" : "node-mongo3-3:27017", "mongoVersion" : "3.2.17", "ping" : ISODate("2021-09-04T15:53:14.436Z"), "up" : NumberLong(17679988), "waiting" : false } { "_id" : "node-mongo2-3:27017", "mongoVersion" : "3.2.17", "ping" : ISODate("2021-10-25T19:15:04.779Z"), "up" : NumberLong(22098498), "waiting" : true } { "_id" : "node-mongo1-3:27017", "mongoVersion" : "3.2.17", "ping" : ISODate("2021-08-31T15:28:27.695Z"), "up" : NumberLong(17330247), "waiting" : false } rs.printReplicationInfo() configured oplog size: 9775.30859375MB log length start to end: 15014162secs (4170.6hrs) oplog first event time: Wed May 05 2021 23:42:53 GMT+0000 (GMT) oplog last event time: Tue Oct 26 2021 18:18:55 GMT+0000 (GMT) now: Tue Oct 26 2021 18:18:55 GMT+0000 (GMT) CR2ConfigRepSet_Prod:SECONDARY> rs.printSlaveReplicationInfo() source: node-mongo3-3.mydomain.com:27019 syncedTo: Tue Oct 26 2021 18:19:11 GMT+0000 (GMT) 0 secs (0 hrs) behind the primary source: node-mongo1-3.mydomain.com:27019 syncedTo: Tue Oct 26 2021 18:19:11 GMT+0000 (GMT) 0 secs (0 hrs) behind the primary source: node-mongo3-504.mydomain.com:27019 syncedTo: Tue Oct 26 2021 18:19:11 GMT+0000 (GMT) 0 secs (0 hrs) behind the primary
解答
该现象不属于正常运行状态下的预期表现,但不会影响本次Balancer停止操作的生效,具体原因和处理方案如下:
- 告警的核心原因是配置服务器的
config.mongos集合中留存了已经下线的旧mongos节点记录。MongoDB停止Balancer时,会等待所有注册在案的mongos节点确认收到配置更新,已经离线的节点无法上报新的ping时间,所以触发了"可能存在迁移或节点宕机"的告警。 - 从给出的副本集同步状态可以确认,当前分片集群的存储层完全健康:所有副本节点同步延迟为0,没有数据同步异常。
- 日志末尾已经输出
Balancer is not running now,说明Balancer已经成功停止,本次备份操作可以正常执行,不会受告警影响。
修复方案
执行以下操作清除失效mongos节点记录,后续操作就不会再出现同类告警:
- 连接到任意可用mongos节点,切换到config数据库:
use config - 执行删除命令移除失效记录:
db.mongos.remove({_id: "node-mongo3-3:27017"}) db.mongos.remove({_id: "node-mongo1-3:27017"})
内容的提问来源于stack exchange,提问作者Sanjay
相关产品推荐
相关产品推荐

