You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

生产环境MongoDB三节点副本集长时间压测后单个节点变为(not reachable/healthy)的原因排查请求

Alright, let's dig into your MongoDB replica set issue—this is a common scenario when running multiple instances on a single server under sustained load, so let's break down the likely root causes and how to troubleshoot them.

Key Observations First

From the replica set status you shared, the critical clues are:

  • The failed node (<host>:27018) shows Connection refused in the last heartbeat message, and its last heartbeat receive was hours before it was marked unhealthy
  • Short-term load works fine, but long-term pressure testing triggers the failure
  • Rejoining the node works in dev, so the node itself isn't permanently broken

This points strongly to runtime resource exhaustion or configuration limits that only hit under sustained load, since all three instances are competing for the same server resources.

1. System Resource Depletion (Most Likely Root Cause)

Running three mongod instances on a single server means they're fighting for CPU, memory, disk IO, and file descriptors. Under long-term load, one instance will hit a resource wall first.

Troubleshooting Steps:

  • Check the mongod logs for the failed node: Open the log file for port 27018 and search for keywords like:
    • terminated or killed (indicates the process was stopped by the OS)
    • out of memory or OOM (signs the instance ran out of RAM)
    • too many open files (hit the file descriptor limit)
  • Check system logs: Look in /var/log/syslog (Debian/Ubuntu) or /var/log/messages (RHEL/CentOS) for entries about the OOM Killer—you'll see lines like killed process <pid> (mongod) if the OS killed the instance to free up memory.
  • Monitor resources during load: Use tools like htop, vmstat, or iostat while running your JMeter test:
    • Watch for CPU usage hitting 100% (mongod is CPU-intensive under write load)
    • Check if memory is fully utilized with high swap usage (swap thrashing makes processes unresponsive)
    • Look for high disk IO wait (%iowait in iostat)—three instances writing to the same disk will create a bottleneck, slowing down heartbeats and replication.

2. MongoDB Configuration Misalignment

Even though you said configs are similar, small differences or misconfigured resource limits could cause one instance to fail first.

Troubleshooting Steps:

  • Compare all three mongod.conf files: Pay close attention to:
    • storage.wiredTiger.engineConfig.cacheSizeGB: The sum of cache limits across all three instances should not exceed ~60% of your server's available RAM (leave room for the OS and other processes). If you have 16GB RAM and each instance is set to 6GB, that's 18GB total—guaranteed to trigger OOM.
    • net.maxIncomingConnections: A low limit could cause the instance to reject new connections (including heartbeats) under load.
    • systemLog.verbosity: Higher log levels increase disk IO, which could push an instance over the edge.
  • Check process limits: Verify that the user running mongod has enough file descriptors and process limits set (use ulimit -a for the mongod user—look for open files and max user processes).

3. Replication Sync Overload

Under heavy write load, secondary nodes have to replicate all writes from the primary. If a secondary can't keep up due to disk IO constraints, it might become unresponsive, leading to heartbeat timeouts.

Troubleshooting Steps:

  • Check replication lag: Before the node fails, run rs.status() and look at optimeDate for each node—if the failed node's optimeDate is far behind the primary, it's struggling to sync.
  • Check disk performance: Even on the same server, disk partitions might have different speeds (unlikely, but possible). Use dd if=/dev/zero of=testfile bs=1G count=1 oflag=direct to test write speed on each instance's dbpath directory.

Temporary Fixes & Long-Term Solutions

  • Temporary relief:
    • Reduce the cacheSizeGB for each instance so the total is within safe limits.
    • Pause long-running tests if resources spike, or use bulk inserts instead of single-document writes to reduce load.
  • Long-term fixes:
    • Don't run multiple replica set nodes on a single server—this defeats the purpose of high availability and creates a single point of resource failure. Spread nodes across separate servers.
    • Add monitoring (e.g., Prometheus + Grafana, or MongoDB's built-in tools) to track resource usage, replication lag, and node health in real time.
    • Optimize your test workload: Use batch inserts, avoid unnecessary indexes during testing, and ensure your write pattern matches production as closely as possible.

内容的提问来源于stack exchange,提问作者dwijekom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 02:09:07