Google Cloud Run频繁重启实例致MongoDB连接超限求助
Hey there, sorry to hear you've been stuck with this frustrating issue for a whole year—let's break this down and get you sorted out.
First: Does Cloud Run send signals to Node.js processes?
Yes, Cloud Run does send a SIGTERM signal to your container's main process when it's about to terminate it (either for scaling down, a new deployment, or health-related restarts). The catch? You have 30 seconds to handle this signal before Cloud Run sends an uncatchable SIGKILL which will force-terminate the process immediately.
Looking at your current graceful shutdown code, there's a critical issue: the exit event in Node.js is synchronous—you can't use async/await there, because Node won't wait for the Promise to resolve before exiting. That's why your mongoose.connection.close() call never completes, and the connection isn't properly cleaned up.
Fixing the Graceful Shutdown Logic
Here's a revised version that properly handles Cloud Run's signals and waits for MongoDB connections to close:
const gracefulShutdown = async () => { console.log('Initiating graceful shutdown...'); try { await mongoose.connection.close(); console.log('✅ MONGODB CONNECTION CLOSED SUCCESSFULLY'); process.exit(0); // Exit cleanly once connections are closed } catch (err) { console.error('❌ Error closing MongoDB connection:', err); process.exit(1); // Exit with error code if cleanup fails } }; // Handle SIGTERM (sent by Cloud Run on container termination) process.on('SIGTERM', async () => { await gracefulShutdown(); }); // Handle SIGINT (for local testing, e.g., Ctrl+C) process.on('SIGINT', async () => { await gracefulShutdown(); }); // Handle uncaught exceptions to avoid dirty exits process.on('uncaughtException', async (err) => { console.error('Uncaught exception:', err); await gracefulShutdown(); }); // Skip the 'exit' event handler—async code won't work here!
How to Identify Container Restarts or New Instance Creations in Cloud Run
You can track instances and restarts using these methods:
- Environment Variables: Cloud Run injects unique identifiers into each instance:
K_SERVICE: Your service nameK_REVISION: The current revision of your serviceK_CONFIGURATION: The configuration name for your revision
- Instance Metadata: You can fetch a unique instance ID via the metadata server (run this once at app startup):
const fetchInstanceID = async () => { const response = await fetch('http://metadata.google.internal/computeMetadata/v1/instance/id', { headers: { 'Metadata-Flavor': 'Google' } }); const instanceId = await response.text(); console.log(`Current instance ID: ${instanceId}`); return instanceId; }; - Cloud Logging: Look for log entries with:
Started containerfor new instance launchesTerminated containerfor instance shutdowns- Log entries tagged with the instance ID (you'll see this in the log metadata)
Fixing the MongoDB Connection Flood
The root cause of your connection limit hitting 1500 is almost certainly orphaned connections from terminated Cloud Run instances that never got closed. Combine the fixed graceful shutdown with these connection pool optimizations:
exports.connect = () => { // Avoid creating duplicate connections if one already exists if (mongoose.connection.readyState === 1) { console.log('ℹ️ Reusing existing MongoDB connection'); return mongoose.connection; } try { mongoose.connect(MONGO.URI, { useCreateIndex: true, keepAlive: true, keepAliveInitialDelay: 300000, // 5 minutes to prevent idle connection drops useNewUrlParser: true, useUnifiedTopology: true, useFindAndModify: false, // Configure connection pool to match your scale poolSize: 8, // Adjust based on your max instance count (total connections = instances × poolSize) maxPoolSize: 8, minPoolSize: 2, serverSelectionTimeoutMS: 5000, // Fail fast if connection can't be established }) .then(() => { console.log('✅ DB Connected successfully'); }) .catch((err) => { console.error('❌ DB Connection Error:', err); }); return mongoose.connection; } catch (err) { console.error('#### Fatal Error Connecting DB', err); console.log('Mongo URI:', MONGO.URI); throw err; } };
Additional Checks to Debug
- Check Cloud Run Instance Restart Triggers: Go to your Cloud Run service's "Logs" tab and look for why instances are restarting—are you hitting memory/cpu limits? Are there health check failures? If Cloud Run is force-killing instances (e.g., due to memory exhaustion), it may send
SIGKILLdirectly without waiting forSIGTERM. - Monitor MongoDB Connections: Use MongoDB's built-in tools like
db.serverStatus().connectionsto track active vs. idle connections and confirm if orphaned connections are being cleaned up after implementing the fixes. - Limit Cloud Run Instance Count: If you're scaling to many instances, adjust the
poolSizeso total connections (instances × poolSize) stay well below MongoDB's 1500 limit.
内容的提问来源于stack exchange,提问作者CaptainZero

