You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MongoDB一主两从副本集随机自动删除集合问题求助

Troubleshooting Random Collection Deletions in Your MongoDB Replica Set

Hey Rilwan, let's dig into this frustrating issue you're dealing with—random collections vanishing, paired with that confusing log entry about reevaluating sync sources. Let's break down what's going on and how to fix it.

First, Let's Decode That Log Entry

The two key red flags in your log are:

  • Every node shows term: -1
  • The sync source's timestamp seems backwards (Mar 31 being "behind" Mar 21—this is likely a timezone or log formatting quirk, but the core issue is the massive 30+ second gap in opTimes between nodes)

Why term: -1 Is a Critical Problem

In MongoDB replica sets, term tracks election cycles—only uninitialized nodes or ones that haven't properly joined the replica set will show -1. If all your nodes have this value, your "replica set" isn't actually a coordinated cluster—each node is probably running as a standalone instance, not syncing properly with the others.

This misconfiguration is almost certainly tied to your disappearing collections. When nodes think they're part of a replica set but aren't properly synced (thanks to no valid election state), the sync logic gets confused. It might incorrectly try to "align" data by deleting collections that exist on other nodes, or overwrite data with outdated states when switching sync sources.

The Sync Source Gap: What It Means

That 30+ second gap between your sync source and mongodb03 suggests either:

  • Your nodes have wildly out-of-sync system clocks (MongoDB needs clocks to be within 10 seconds of each other to sync properly)
  • There's massive network latency between nodes, breaking the sync pipeline
  • One of your nodes is stuck in an outdated state (like it hasn't synced in days, hence the Mar 21 timestamp)

Step-by-Step Fixes

Let's walk through how to diagnose and resolve this:

1. Verify Your Replica Set Is Actually Working

Log into your supposed primary node and run this command to check the cluster state:

rs.status()

Look for these details:

  • In the members array, each node should have a stateStr of PRIMARY or SECONDARY (if you see STARTUP or UNKNOWN, the replica set never initialized properly)
  • Each node's optime.term should be a positive integer (not -1)

If the replica set isn't initialized, fix that first:

# On the intended primary node
rs.initiate()
# Add your secondary nodes
rs.add("mongodb02:27017")
rs.add("mongodb03:27017")

2. Fix Clock Sync Across Nodes

MongoDB relies on synchronized system clocks to correctly order operations. Even a small drift can break sync. On each node, run:

date

Make sure all nodes show the exact same time and timezone. If they don't, set up NTP (Network Time Protocol) on every node to automatically sync clocks—this is non-negotiable for replica set stability.

3. Correct Sync Source Configuration

Sometimes secondary nodes pick bad sync sources (like other secondaries instead of the primary), leading to delays and sync errors. Check your replication status with:

rs.printSlaveReplicationInfo()
rs.printReplicationInfo()

If a secondary is syncing from another secondary, force it to sync directly from the primary:

# Run this on the problematic secondary node
rs.syncFrom("mongodb01:27017") # Replace with your primary's address

4. Rule Out Accidental Human/Application Errors

Even though you said collections are "automatically" deleted, it's worth double-checking:

  • Do any application scripts or cron jobs run dropCollection commands?
  • Could someone have accidentally run deletions on the wrong node?

If you have auditing enabled, check those logs. If not, temporarily enable profiling to track collection drops:

db.setProfilingLevel(2, { slowms: 0 })
# Later, check the profile logs with
db.system.profile.find({ op: "drop" })
# Don't forget to disable profiling when done
db.setProfilingLevel(0)

Recovering Deleted Collections

If you have a recent backup, here's how to restore:

  1. Stop replication on your secondary nodes to avoid overwriting data:
    rs.stopSecondaryReplication()
    
  2. Restore the deleted collection from backup to your primary node.
  3. Restart replication on the secondaries:
    rs.startSecondaryReplication()
    

内容的提问来源于stack exchange,提问作者Rilwan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:38:49