GOMAXPROCS从1到4性能线性提升,4到8趋于平缓的原因问询
Problem Description
I'm using an 8-core Mac equipped with a 2.8 GHz Intel Core i7 processor, and I've confirmed the core count via fmt.Println(runtime.NumCPU()). I implemented a simple worker pool model to handle CPU-intensive tasks concurrently, aiming to understand how performance scales when allocating more cores to Go.
Here's my implementation code:
func Run(poolSize int, workSize int, loopSize int, maxCores int) { runtime.GOMAXPROCS(maxCores) var wg sync.WaitGroup wg.Add(poolSize) defer wg.Wait() // Channel for sending pending requests to the worker pool workStream := make(chan int) // cpuIntensiveWork simulates a CPU-bound task var cpuIntensiveWork = func(input int) { res := input for i := 0; i < loopSize; i++ { res = res + i } } // worker is the processing function launched by the pool worker := func(wg *sync.WaitGroup, workStream chan int, id int) { defer wg.Done() for req := range workStream { cpuIntensiveWork(req) } } // Launch worker goroutines for i := 0; i < poolSize; i++ { go worker(&wg, workStream, i) } // Send tasks to workStream and close the channel once done for workItemNo := 0; workItemNo < workSize; workItemNo++ { workStream <- workItemNo } close(workStream) }
And the benchmark code:
var numberOfWorkers = 100 var numberOfRequests = 1000 var loopSize = 100000 func Benchmark_1Core(b *testing.B) { for i := 0; i < b.N; i++ { Run(numberOfWorkers, numberOfRequests, loopSize, 1) } } func Benchmark_2Cores(b *testing.B) { for i := 0; i < b.N; i++ { Run(numberOfWorkers, numberOfRequests, loopSize, 2) } } func Benchmark_4Cores(b *testing.B) { for i := 0; i < b.N; i++ { Run(numberOfWorkers, numberOfRequests, loopSize, 4) } } func Benchmark_8Cores(b *testing.B) { for i := 0; i < b.N; i++ { Run(numberOfWorkers, numberOfRequests, loopSize, 8) } }
After running the benchmarks, I noticed that performance scales almost linearly when increasing GOMAXPROCS from 1 to 2, and 2 to 4. However, the performance gain is very limited when going from 4 to 8 cores. Is this expected behavior? If so, what's the reason behind it?
Answer
This is completely expected behavior, and there are a few key reasons behind it:
Intel Hyper-Threading Limitations: Your 2.8 GHz Intel Core i7 uses Hyper-Threading technology—those 8 "cores" are actually 4 physical cores plus 4 logical cores (threads) that share the physical core's execution units (like ALUs, caches, etc.). For CPU-intensive tasks, logical cores don't deliver the same performance boost as physical cores. Typically, Hyper-Threading only provides a 10-30% performance gain per logical core, not the 100% linear improvement you get from adding physical cores. So moving from 4 physical cores (
GOMAXPROCS=4) to 8 logical cores (GOMAXPROCS=8) can't replicate the linear scaling you saw with physical cores.Resource Contention for CPU-Bound Tasks: Since your tasks are purely CPU-intensive, once you saturate the 4 physical cores, adding more logical cores means goroutines will compete for shared resources on the same physical core. This can lead to cache invalidations and increased context-switching overhead, which can offset some of the gains from Hyper-Threading, resulting in minimal overall performance improvement.
Increased Scheduling Overhead: When setting
GOMAXPROCS=8, the Go scheduler has to manage more OS threads (M in Go's M-P-G model). This adds a small but measurable overhead, especially when the CPU is already near full utilization, further limiting the performance gain.
In short, your test perfectly demonstrates the limitations of Hyper-Threading in CPU-bound scenarios—physical cores deliver reliable linear scaling, while logical cores offer much more modest gains.
内容的提问来源于stack exchange,提问作者Picci

