Go并发性能异常:高QPS下执行时长攀升但CPU利用率低
Hey ChrisDave, let's dig into this tricky issue you're facing—this is a super common gotcha with Go concurrency when CPU isn't the bottleneck, so let's break down the likely culprits and how to fix them.
核心问题:低CPU使用率但响应时间飙升
When your CPU is sitting at 3-6% but request latency keeps climbing, it means your goroutines aren't doing CPU work—they're blocking on something else. Here are the most likely causes:
1. Channel阻塞(缓冲不足或无缓冲)
If you're using channels to coordinate goroutines, a small or unbuffered channel can become a bottleneck fast. For example, if you're sending every request to a worker channel with a tiny buffer, hundreds of goroutines will get stuck waiting to send data instead of processing requests.
Bad example:
// 缓冲仅10的worker channel,QPS1200时瞬间被占满 var workerChan = make(chan Request, 10) func handleRequest(w http.ResponseWriter, r *http.Request) { req := parseRequest(r) workerChan <- req // 大量goroutine在这里阻塞,CPU无事可做 }
Fix: Use a worker pool with a buffer size matching your expected concurrency, or dynamically adjust worker count based on load.
2. IO阻塞(最常见的元凶)
If your program makes database calls, HTTP requests, or file operations without proper connection pooling, goroutines will spend most of their time waiting for IO instead of using CPU. For example:
- Not setting
MaxOpenConnson yourdatabase/sqlconnection pool, leading to goroutines waiting for available DB connections - Creating a new HTTP client for every request instead of reusing a client with a
Transportthat has enough idle connections
Even if CPU is free, blocked goroutines pile up and increase overall latency.
3. 不当的同步原语使用
If you're using sync.Mutex or sync.RWMutex with overly broad scope, or relying on global locks for shared resources, goroutines will queue up waiting for the lock instead of processing work. For example, a global lock around a cache access that every request hits will turn your concurrent program into a single-threaded one—CPU stays idle, but latency skyrockets.
4. Goroutine泄漏或调度开销
If each request spawns a goroutine that never exits (e.g., waiting on a channel that's never closed, or stuck in an infinite loop), the number of goroutines will grow over time. Go's scheduler can handle thousands of goroutines, but as the count climbs into tens of thousands, scheduling overhead increases—even if each goroutine is blocked, the scheduler has more work to do, leading to higher latency.
排查步骤
Here's how to pinpoint the issue:
- Use
go tool pproffor blocking/mutex profiles: These profiles show exactly where your goroutines are getting stuck. Run:# 采集30秒的阻塞profile go tool pprof http://localhost:6060/debug/pprof/block?seconds=30 # 采集mutex竞争profile go tool pprof http://localhost:6060/debug/pprof/mutex?seconds=30 - Monitor goroutine count: Add a metric or log using
runtime.NumGoroutine()—if it's steadily increasing, you've got a leak. - Check connection pool settings: Verify DB, Redis, and HTTP client pools have enough max connections to handle your QPS.
- Validate channel buffer sizes: Ensure worker channels have enough buffer to handle peak load, or use a fixed worker pool to limit concurrent processing.
总结
Your problem isn't about CPU power—it's about goroutines being blocked on resources outside the CPU. Focus on finding those blocking points with profiling tools, and adjust your concurrency model, connection pools, or channel usage to eliminate the bottlenecks.
内容的提问来源于stack exchange,提问作者ChrisDave

