Linux 4.4.1众核系统XPS配置后未调用get_xps_queue()问题求助
Hey there, let's break down how to get that strict 1:1 core-to-transmit queue mapping working for your setup. Since you're already running with XPS enabled but relying solely on skb_tx_hash for queue selection, here are actionable approaches tailored to Linux 4.4.1:
Approach 1: Tweak skb_tx_hash for Direct CPU-to-Queue Mapping
The simplest kernel code change is to modify the skb_tx_hash function (in net/core/dev.c) to directly map the current CPU core to a TX queue, instead of using packet hash. This guarantees each core's traffic goes to its dedicated queue—assuming you’ve already set the number of TX queues equal to your core count n.
Example Code Modification:
Replace the existing hash-based logic with CPU ID mapping:
u16 skb_tx_hash(const struct net_device *dev, const struct sk_buff *skb) { /* Map current CPU to TX queue 1:1 (ensure dev->real_num_tx_queues == n) */ return raw_smp_processor_id() % dev->real_num_tx_queues; }
- Use
raw_smp_processor_id()instead ofsmp_processor_id()to safely get the current CPU even in softirq contexts (common for packet transmission). - First confirm your TX queue count matches core count with
ethtool -l <iface>; adjust it if needed withethtool -L <iface> tx n.
Approach 2: Enforce XPS Mapping via Sysfs (No Kernel Code Changes)
Since you have CONFIG_XPS=y enabled, you can leverage XPS's per-queue CPU masks to lock each core's traffic to a specific TX queue. This works if your network stack is invoking get_xps_queue() (if not, you’ll need a small code tweak to ensure it’s called in __dev_queue_xmit).
Step-by-Step Sysfs Configuration:
Assume your network interface is eth0 and you have n cores/queues:
# Loop through each TX queue and bind it to the corresponding core for i in $(seq 0 $((n-1))); do # Set xps_cpus to only the i-th core (bitmask: 1 << i) echo 0x$(printf "%x" $((1 << i))) > /sys/class/net/eth0/queues/tx-$i/xps_cpus done
To ensure get_xps_queue() is used for queue selection, modify __dev_queue_xmit in net/core/dev.c to check for XPS queues first:
// Replace the existing queue_index assignment with: struct netdev_queue *xps_queue = get_xps_queue(dev, skb); if (xps_queue) { queue_index = xps_queue->index; } else { queue_index = skb_tx_hash(dev, skb); }
Approach 3: Modify NIC Driver Queue Selection
If you’re using a specific network card driver (e.g., ixgbe, e1000e), you can directly adjust the driver’s queue selection logic in its ndo_start_xmit handler. This gives you fine-grained control without touching core network code.
Example Driver Tweak:
In the driver’s transmit function, replace the queue selection with:
static netdev_tx_t my_driver_start_xmit(struct sk_buff *skb, struct net_device *dev) { int queue_idx = raw_smp_processor_id() % dev->real_num_tx_queues; // Rest of the transmit logic using queue_idx... }
Bonus: Optimize Interrupt Affinity
To complete the 1:1 mapping, bind each TX queue’s interrupt to its corresponding core. This reduces cross-core traffic when processing completed transmits:
# Find TX queue interrupts (adjust interface name as needed) grep "<iface>-Tx" /proc/interrupts # Bind each interrupt to its matching core (e.g., interrupt 123 to core 0) echo 0 > /proc/irq/123/smp_affinity_list
Choose the approach that fits your constraints—sysfs-based XPS is best if you want to avoid kernel patches, while modifying skb_tx_hash is the most direct for strict mapping.
内容的提问来源于stack exchange,提问作者Ramneek

