MongoDB 3.0.14副本集Secondary节点Global锁等待过长问题求助
Let’s break down your problem using the query logs, explain output, and currentOp snapshot you’ve shared:
1. Why are Global lock wait times so high?
Your logs show queries waiting 4.5s to 21s to acquire a Global intent-shared (r) lock on a Secondary node. Even with WiredTiger’s document-level locking, Global locks can become a bottleneck in these scenarios:
- Oplog application overhead: Secondaries sync with the Primary by applying oplog entries. In MongoDB 3.0, oplog application (especially for bulk writes or large transaction batches) can occasionally hold Global locks for extended periods, blocking incoming read queries.
- Background system operations: WiredTiger’s periodic checkpointing, log rotation, or replica set maintenance tasks (like syncing new nodes) may briefly hold Global locks. High production load can stretch these normally short waits into noticeable delays.
- Ephemeral long-running operations: If the lock-holding operation completed right before you ran
db.currentOp(), it won’t show up in the results. For example, a large oplog apply batch or a hidden maintenance task could have blocked your reads before you checked. - MongoDB 3.0 WiredTiger limitations: The 3.0 series was the first to integrate WiredTiger, and later versions (3.2+) fixed several lock contention bugs. Your older version may be prone to these edge cases.
Your currentOp snapshot confirms this: a getmore operation is stuck waiting for a Global r lock, meaning a competing operation held the lock long enough to block subsequent read requests.
2. How to capture operations holding Global R/W locks when db.currentOp() doesn’t show anything?
db.currentOp() only shows active operations—if the lock-holding task finished before you ran the command, you won’t see it. Try these approaches to catch the culprit:
- Enable verbose lock logging: Temporarily set
logLevel=5(or dynamically rundb.setLogLevel(5, "lock")) to log detailed lock acquisition/release events. This will show exactly which operations hold Global locks and their duration. - Use the system profiler: Enable full profiling (level 2) with
db.setProfilingLevel(2)to log all operations. Filter the profiler logs to find tasks holding heavy Global locks:
Focus on entries with highdb.system.profile.find({ "locks.Global": { $in: ["X", "S"] } })millisvalues—these are the operations likely causing waits. - Real-time monitoring tools:
mongostat: Track lock acquisition rates and wait times (look atqr,qw,ar,awmetrics to spot contention).mongotop: Identify which collections/operations are consuming the most resources, helping pinpoint lock-heavy tasks.
- Check oplog sync metrics: On the Secondary, run
db.serverStatus().replto view oplog apply lag and batch sizes. Large batches or persistent lag can signal prolonged lock usage during sync.
3. Why does the query need to acquire the Global intent shared lock 14 times?
The acquireCount correlates directly with how many times the query yields and re-acquires locks:
- When a query yields (to let other operations run), it releases all held locks. Once it resumes, it must re-acquire the necessary locks to continue execution.
- Your first query has
numYields:6—each yield adds two lock operations (release + re-acquire) plus the initial lock acquisition. The math roughly adds up:1 (initial) + 2*6 (yield/re-acquire) = 13, with the extra count coming from internal cursor operations (likegetmorecalls or index traversal stages). - Your
explainoutput showssaveState:2andrestoreState:2, which correspond to 2 yields during the low-load explain run. In production, higher contention leads to more yields, hence higheracquireCount.
WiredTiger uses intent locks to signal that a transaction plans to lock lower-level resources (like documents or collections). Each time the query resumes after a yield, it needs to re-establish these intent locks at the Global level.
Additional Recommendations
- Upgrade MongoDB: Version 3.0.14 is end-of-life and has known WiredTiger lock contention bugs. Upgrading to a supported version (e.g., 3.6 or later) will give you access to improved locking logic and stability fixes.
- Optimize read queries: While your
explainshows efficient index usage, reducing result set size (e.g., using projection to limit returned fields) can decrease yield frequency and lock acquisition overhead. - Monitor replica set health: Keep an eye on oplog lag and Secondary sync status to ensure oplog application isn’t causing unexpected lock contention.
内容的提问来源于stack exchange,提问作者Abhishek Kumar

