Google Cloud Run低负载下异常扩容问题咨询
Why Your Cloud Run Service is Over-Scaling with Low Traffic (And How to Fix It)
Hey there, let’s break down why your low-traffic Express/Node.js Cloud Run service is scaling up to 4 instances and sticking at 2, even when CPU usage is nearly zero. Then we’ll cover actionable fixes to get it aligned with your actual needs.
Possible Reasons for Unwanted Scaling
- Cold start & request queueing compensation: Cloud Run prioritizes reducing request latency. If your app has even a slight cold start delay (common with Node.js/Express), the system might spin up extra instances preemptively when it detects incoming requests—even if those requests are low-volume. Once instances are up, they don’t shut down immediately (default idle timeout is 15 minutes), so you’ll see lingering instances even after traffic drops.
- Concurrency configuration mismatch: Your max concurrency is set to 40, which means each instance can handle up to 40 requests at once. But if some requests take longer than expected (e.g., a slow database call, external API fetch), Cloud Run might interpret that as the instance being overwhelmed and spin up new instances to take on pending requests—even if the total number of requests is tiny.
- Idle instance retention: Even with
min-instancesset to 0, Cloud Run may keep idle instances around temporarily as a "warm pool" to avoid future cold starts. This is especially true if your service has had recent traffic spikes; the system might hold onto 1-2 instances longer than you expect. - Hidden traffic: Don’t overlook health checks, liveness probes, or even automated crawlers. These can count towards your request volume, triggering scaling even when user traffic is minimal. Check your Cloud Logging to see if unexpected requests are hitting your service.
Optimization Fixes to Try
- Tune max concurrency down: Lower your max concurrency from 40 to a smaller number (like 5-10). This tells Cloud Run that each instance should handle fewer requests at once, reducing the likelihood of it spinning up new instances for small bursts of traffic. Since your load is extremely low, a lower concurrency limit will keep instances busy enough without needing to scale.
- Set min-instances to 1 (instead of 0): Counterintuitive as it sounds, keeping one warm instance can prevent unnecessary scaling. When
min-instancesis 0, every incoming request triggers a cold start, and Cloud Run might spin up extra instances to compensate for the delay. A single warm instance eliminates cold starts and keeps your scaling behavior predictable. - Shorten instance idle timeout: By default, idle instances stay running for 15 minutes. You can reduce this via the Cloud Run console or CLI with
gcloud run services update [SERVICE_NAME] --max-instance-idle-time=1m(set to 1 minute, for example). This ensures idle instances shut down quickly once traffic stops, preventing lingering instances from inflating your count. - Limit max instances to 1: If you’re confident a single instance can handle all your occasional concurrent requests, set
max-instancesto 1. This hard caps scaling, so Cloud Run can’t spin up more than one instance no matter what. Just make sure to test this with your peak expected traffic to confirm it doesn’t cause latency issues. - Audit request logs: Dig into Cloud Logging to identify any unexpected traffic (health checks, bots, etc.). If health checks are contributing to scaling, you can adjust their frequency or configure them to avoid triggering instance creation.
内容的提问来源于stack exchange,提问作者Bruce C
相关产品推荐
相关产品推荐

