Azure Standard_L64s_v3实例单核心CPU下Netperf TCP_RR性能提升原因咨询
Let me break down the technical reasons behind this counterintuitive result—this is actually a common scenario with latency-sensitive workloads like TCP_RR (Request-Response):
1. Your TCP_RR test is single-threaded, so extra cores add overhead instead of value
Your 64-core test uses -T 1,1 to pin the Netperf client to a single core, meaning the other 63 cores are idle but still part of the system's scheduler context. Here's the catch:
- The kernel has to manage scheduler ticks, process queues, and potential thread migrations across all 64 cores, even for your single Netperf thread. This introduces unnecessary scheduling overhead that eats into latency-sensitive operations.
- Even with pinning, subtle kernel-level activities can sometimes trigger cross-core thread migration, which flushes the CPU cache and forces the thread to reload its working set (socket buffers, connection state) from slower memory.
2. Cache locality eliminates coherence overhead
On a 64-core system, the cache hierarchy has to maintain coherence across all cores. Background processes, interrupt handlers, and system daemons running on other cores generate constant cache coherence traffic, adding latency to your Netperf thread's memory accesses.
When you limit to 1 core:
- All memory accesses stay within the single core's L1/L2 cache (and shared L3 if applicable, with no competing cores), wiping out cross-core cache coherence overhead.
- The Netperf thread's critical data (socket buffers, request/response state) stays "hot" in the cache, cutting down memory access latency drastically—this is make-or-break for TCP_RR, where every cycle depends on fast data retrieval.
3. Interrupt handling stays local to the test thread
On multi-core systems, network interrupts are often spread across cores via IRQ balancing to distribute load. For TCP_RR, this means the interrupt handler (processing incoming network data) might run on a different core than your Netperf thread. The thread has to wait for data to be passed between cores, adding measurable latency to each request/response cycle.
With 1 core:
- All network interrupts and the Netperf thread run on the same core. There's no inter-core communication delay for passing network data from the interrupt handler to the application thread, shaving off latency at every step.
4. Reduced kernel housekeeping overhead
A 64-core system runs more kernel threads and housekeeping processes (like per-core kworkers, scheduler helpers, and memory management daemons) to manage all those cores. These consume small amounts of CPU time and memory bandwidth, which adds up for a latency-sensitive workload like TCP_RR.
On 1 core, the kernel has far less housekeeping to handle, so nearly all CPU resources and memory bandwidth are dedicated exclusively to your Netperf thread.
Quick recap of your test context
64-core client test:
- Server command:
netserver -4 -v -d -N -p- Client command:
netperf -H <目标IP> -p <端口> -l 60 -T 1,1 -t TCP_RR- Result: 9147.83 transactions per second (TPS)
1-core client test (kernel cmdline:
maxcpus=1 nr_cpus=1):
- Client command:
netperf -H <目标IP> -p <端口> -l 60 -t TCP_RR- Result: 10183.33 TPS
If you wanted to leverage all 64 cores, you'd need to run multiple parallel Netperf TCP_RR instances (one per core) and sum their TPS—this would show a clear improvement over the single-core single-instance result. The key here is that TCP_RR is latency-bound and single-threaded by nature, so extra cores don't help unless you scale the workload itself.
内容的提问来源于stack exchange,提问作者user1994587

