HAProxy多进程负载不均求助:nbproc 10配置下进程间RPS差异过大
Hey there, let's break down why your HAProxy setup with nbproc 10 is seeing such lopsided RPS across processes on that Xeon E5-1650 v3. I’ve tackled similar issues before, so here are the key areas to check and fixes to implement:
1. Fix CPU Affinity (Critical First Step)
Your Xeon E5-1650 v3 has 6 physical cores with hyper-threading enabled, giving you 12 logical CPUs. When you run 10 HAProxy processes without binding them to specific cores, the kernel’s scheduler can bounce processes around, leading to uneven load.
- First, map out your CPU topology with
lscpuorcat /proc/cpuinfoto get logical CPU IDs (usually 0-11 for this chip). - Add
cpu-mapdirectives to your HAProxy config to pin each process to a dedicated logical CPU. Example:
(Left number = HAProxy process ID, right number = logical CPU ID)nbproc 10 cpu-map 1 0 cpu-map 2 1 cpu-map 3 2 cpu-map 4 3 cpu-map 5 4 cpu-map 6 5 cpu-map 7 6 cpu-map 8 7 cpu-map 9 8 cpu-map 10 9 - This eliminates scheduler thrashing and ensures each process has consistent access to CPU resources.
2. Adjust Connection Distribution Logic
HAProxy’s main process distributes new connections to child processes, and default behavior can lead to imbalance if you have skewed source IPs or long-lived connections:
- If using HTTP mode: Enable
option http-server-closeto reduce persistent client connections that might stick to specific processes. This forces shorter-lived connections, making distribution more even. - Check connection balancing: For TCP frontends, ensure you’re using a balanced distribution strategy. If you’re relying on source IP hashing (
balance source), switch tobalance roundrobinorbalance leastconnfor more even load spread—especially if your traffic has a small set of repeat source IPs. - Verify current load: Use HAProxy’s stats socket to check per-process connections and RPS:
Look atecho "show stat" | socat stdio /var/run/haproxy.sockreq_rate(RPS) andconn_cur(current connections) columns for each process to confirm if imbalance ties to long-lived connections.
3. Align Network Interrupt Affinity
Network card interrupts can monopolize certain CPUs, starving nearby HAProxy processes. Fix this by binding NIC interrupts to the same CPUs your HAProxy processes are using:
- Find your NIC’s IRQ number with
cat /proc/interrupts(look for your interface name, e.g.,eth0). - Bind the IRQ to a subset of your HAProxy CPUs. For example, if your NIC has multiple queues, split them across your mapped CPUs:
# Bind IRQ 123 to CPUs 0,2,4,6,8 (match even-numbered logical cores) echo 0,2,4,6,8 > /proc/irq/123/smp_affinity_list - Avoid using
irqbalanceunless you configure it to prioritize keeping NIC interrupts tied to HAProxy cores—defaultirqbalancecan spread interrupts randomly, worsening imbalance.
4. Check HAProxy Version & System Tuning
- Upgrade HAProxy: Older versions (pre-2.0) had known bugs with
nbprocconnection distribution. Upgrade to a stable LTS release like 2.4.x or 2.6.x to resolve any underlying issues. - Boost Process Priority: Give HAProxy processes higher CPU priority to prevent them from being starved by other system tasks:
renice -n -5 $(pidof haproxy) - Confirm Hyper-Threading: Ensure hyper-threading is enabled (
lscpushould showThread(s) per core: 2). With 6 physical cores, 10 processes are manageable, but disabled hyper-threading would mean overcrowding 6 cores, leading to unavoidable imbalance.
5. Validate After Changes
After applying these fixes, restart HAProxy and monitor per-process stats over 10-15 minutes. You should see RPS and connection counts even out across your 10 processes.
内容的提问来源于stack exchange,提问作者artful

