生产环境MongoDB三节点副本集长时间压测后单个节点变为(not reachable/healthy)的原因排查请求
Alright, let's dig into your MongoDB replica set issue—this is a common scenario when running multiple instances on a single server under sustained load, so let's break down the likely root causes and how to troubleshoot them.
From the replica set status you shared, the critical clues are:
- The failed node (
<host>:27018) showsConnection refusedin the last heartbeat message, and its last heartbeat receive was hours before it was marked unhealthy - Short-term load works fine, but long-term pressure testing triggers the failure
- Rejoining the node works in dev, so the node itself isn't permanently broken
This points strongly to runtime resource exhaustion or configuration limits that only hit under sustained load, since all three instances are competing for the same server resources.
1. System Resource Depletion (Most Likely Root Cause)
Running three mongod instances on a single server means they're fighting for CPU, memory, disk IO, and file descriptors. Under long-term load, one instance will hit a resource wall first.
Troubleshooting Steps:
- Check the mongod logs for the failed node: Open the log file for port 27018 and search for keywords like:
terminatedorkilled(indicates the process was stopped by the OS)out of memoryorOOM(signs the instance ran out of RAM)too many open files(hit the file descriptor limit)
- Check system logs: Look in
/var/log/syslog(Debian/Ubuntu) or/var/log/messages(RHEL/CentOS) for entries about the OOM Killer—you'll see lines likekilled process <pid> (mongod)if the OS killed the instance to free up memory. - Monitor resources during load: Use tools like
htop,vmstat, oriostatwhile running your JMeter test:- Watch for CPU usage hitting 100% (mongod is CPU-intensive under write load)
- Check if memory is fully utilized with high swap usage (swap thrashing makes processes unresponsive)
- Look for high disk IO wait (
%iowaitiniostat)—three instances writing to the same disk will create a bottleneck, slowing down heartbeats and replication.
2. MongoDB Configuration Misalignment
Even though you said configs are similar, small differences or misconfigured resource limits could cause one instance to fail first.
Troubleshooting Steps:
- Compare all three
mongod.conffiles: Pay close attention to:storage.wiredTiger.engineConfig.cacheSizeGB: The sum of cache limits across all three instances should not exceed ~60% of your server's available RAM (leave room for the OS and other processes). If you have 16GB RAM and each instance is set to 6GB, that's 18GB total—guaranteed to trigger OOM.net.maxIncomingConnections: A low limit could cause the instance to reject new connections (including heartbeats) under load.systemLog.verbosity: Higher log levels increase disk IO, which could push an instance over the edge.
- Check process limits: Verify that the user running mongod has enough file descriptors and process limits set (use
ulimit -afor the mongod user—look foropen filesandmax user processes).
3. Replication Sync Overload
Under heavy write load, secondary nodes have to replicate all writes from the primary. If a secondary can't keep up due to disk IO constraints, it might become unresponsive, leading to heartbeat timeouts.
Troubleshooting Steps:
- Check replication lag: Before the node fails, run
rs.status()and look atoptimeDatefor each node—if the failed node'soptimeDateis far behind the primary, it's struggling to sync. - Check disk performance: Even on the same server, disk partitions might have different speeds (unlikely, but possible). Use
dd if=/dev/zero of=testfile bs=1G count=1 oflag=directto test write speed on each instance'sdbpathdirectory.
Temporary Fixes & Long-Term Solutions
- Temporary relief:
- Reduce the
cacheSizeGBfor each instance so the total is within safe limits. - Pause long-running tests if resources spike, or use bulk inserts instead of single-document writes to reduce load.
- Reduce the
- Long-term fixes:
- Don't run multiple replica set nodes on a single server—this defeats the purpose of high availability and creates a single point of resource failure. Spread nodes across separate servers.
- Add monitoring (e.g., Prometheus + Grafana, or MongoDB's built-in tools) to track resource usage, replication lag, and node health in real time.
- Optimize your test workload: Use batch inserts, avoid unnecessary indexes during testing, and ensure your write pattern matches production as closely as possible.
内容的提问来源于stack exchange,提问作者dwijekom

