Node.js原生MongoDB驱动在副本集主节点宕机时连接丢失问题
It sounds like your Node.js app isn't handling replica set failover correctly when the primary goes down. Let's break down the key fixes and checks to get this working smoothly, using the native driver as required.
1. Fix Your Connection String
The most common issue here is an incomplete connection string that doesn't tell the driver about the full replica set.
- Include all nodes: List your primary (Win7), secondary (Pi), and arbiter (Pi) in the URI. This ensures the driver knows where to find other nodes if the primary goes down.
- Add the replica set name: You must specify the
replicaSetparameter so the driver recognizes it's connecting to a replica set, not a standalone instance. - Use accessible IPs/hostnames: Avoid
localhostsince your nodes are on different machines—use public/private IP addresses that all nodes and your app can reach.
Example connection string:
const uri = 'mongodb://192.168.1.100:27017,192.168.1.200:27017,192.168.1.200:27018/your-db-name?replicaSet=your-replica-set-name';
2. Configure Driver Connection Options
Newer versions of the native driver (4.x+) handle reconnection automatically, but you need to set sensible timeouts to avoid hanging connections:
const { MongoClient } = require('mongodb'); const client = new MongoClient(uri, { connectTimeoutMS: 30000, // Allow 30s for initial connection attempts socketTimeoutMS: 30000, // Close stale sockets after 30s maxPoolSize: 10, // Adjust based on your app's needs }); // Auto-retry connection on failure async function connect() { try { await client.connect(); console.log('Connected to replica set successfully'); } catch (err) { console.error('Connection failed—retrying in 5s:', err); setTimeout(connect, 5000); } } connect();
3. Validate Replica Set Configuration
Before blaming the driver, ensure your replica set is set up correctly:
- Check node communication: Verify there are no firewall blocks on port 27017/27018 between your Win7 machine and Pi. Test connectivity with
telnet <ip> <port>ornc -zv <ip> <port>. - Use resolvable addresses: When initializing the replica set, use IPs (not
localhost) so all nodes can find each other. Runrs.status()in the MongoDB shell on any node to confirm all nodes are in the correct state (PRIMARY, SECONDARY, ARBITER). - Ensure quorum: With 3 nodes (1 primary, 1 secondary, 1 arbiter), you have a quorum of 2. When the primary goes down, the secondary and arbiter should elect a new primary within 10-30 seconds. If this isn't happening, check your replica set election settings.
4. Handle Connection Events Gracefully
Listen to the driver's events to react to failover and reconnections:
// Log state changes for nodes client.on('serverDescriptionChanged', (event) => { const oldState = event.previousDescription.state; const newState = event.newDescription.state; console.log(`Node ${event.address} changed state: ${oldState} → ${newState}`); if (newState === 'PRIMARY') { console.log('New primary elected:', event.address); } }); // Reconnect if the connection closes client.on('close', () => { console.log('Connection lost—attempting to reconnect'); setTimeout(connect, 5000); }); // Log errors to debug issues client.on('error', (err) => { console.error('MongoDB driver error:', err); });
5. Add Retry Logic for Writes
During failover, there's a short window where no primary exists. Add retry logic for write operations to handle this:
async function safeInsert(collection, document) { let retries = 3; while (retries > 0) { try { return await collection.insertOne(document); } catch (err) { // Check for primary unavailable errors if (err.code === 10107 || err.code === 13435) { retries--; console.log(`Primary unavailable—retrying in 1s (${retries} left)`); await new Promise(resolve => setTimeout(resolve, 1000)); } else { throw err; // Re-throw non-failover errors } } } throw new Error('Failed to insert after 3 retries'); }
Final Testing
Shut down your primary mongod process, then:
- Run
rs.status()on the Pi to confirm a new primary is elected. - Check your app logs to see if it detects the new primary and resumes operations.
- Test writes/reads to ensure they're routed correctly to the new primary.
Content of the question来源于stack exchange,提问作者potter134

