You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

io_uring用于高扇出TCP写入时性能意外劣化的技术咨询

Great question! What you're seeing isn't expected behavior—io_uring absolutely should outperform a loop of non-blocking writes for high-fanout TCP sends. Let's break down why your current implementation is slower and how to fix it.

1. You're not registering file descriptors (FDs) with io_uring

Right now, each IORING_OP_SEND request passes the raw TCP socket FD to the kernel. For every request, the kernel has to look up the FD to get the underlying file structure—this adds overhead that scales linearly with the number of connections.

Fix: Register your active TCP sockets with io_uring's file table using ring.register_files(). Then, reference sockets by their index in the registered table instead of the raw FD. This skips the per-request FD lookup and drastically reduces kernel-side processing time for batches of sends.

Example snippet for registration:

// Collect active FDs first
let fds: Vec<_> = connections.iter()
    .filter(|c| c.state == State::Active)
    .map(|c| c.socket.as_raw_fd())
    .collect();

// Register once (or update when connections change)
ring.register_files(&fds)?;

// When preparing SQEs, use the index instead of FD
sqe.prep_send(
    fds_index, // instead of raw FD
    message.as_ptr(),
    message.len() as u32,
    MSG_NOSIGNAL,
);

2. You're copying data for every send request

Your current setup copies the message data from user space to kernel space for each send request. With many connections, this cumulative copy overhead adds up—and io_uring doesn't magically eliminate this unless you use zero-copy or registered buffers.

Fix: Use io_uring's registered buffers (IORING_REGISTER_BUFFERS) to avoid redundant copies. Register your message buffer once, then reference it by index in each send SQE. The kernel can directly use the user-space buffer (assuming alignment and permissions are correct), eliminating per-request data copies.

Note: For zero-copy sends (IORING_OP_SEND_ZC), you'll need a kernel version ≥ 5.19, and your buffer must be page-aligned (use mmap or aligned allocation). For older kernels, registered buffers still reduce overhead by letting the kernel cache buffer metadata.

3. You're not using SQPOLL for batch submission efficiency

Your current ring.submit() call is a system call that switches from user to kernel mode. For large batches, this context switch plus kernel-side processing of each SQE adds up.

Fix: Create your io_uring instance with the IORING_SETUP_SQPOLL flag. This spawns a kernel thread that polls the submission queue automatically, so you don't need to call submit() at all—just queue your SQEs and let the kernel handle them asynchronously. This eliminates the per-batch system call overhead entirely.

Example ring creation:

let ring = IoUring::new_with_flags(1024, SetupFlags::SQPOLL)?;

4. Kernel version might be holding you back

io_uring has evolved rapidly—older kernels (pre-5.10) lack key optimizations for batch TCP sends, especially around FD lookup and buffer handling. If you're on a kernel older than 5.10, upgrading to a newer stable release (like 5.15 or 6.x) could immediately improve performance.

5. Check your SQE preparation logic

Make sure your push_send function is efficiently preparing SQEs. Avoid unnecessary memory allocations or copies when setting up each request. For example, reuse SQE structures where possible, and ensure you're setting flags like MSG_NOSIGNAL to avoid unnecessary signal handling overhead.

Putting it all together

When you combine FD registration, registered buffers, and SQPOLL, you should see a massive drop in submission latency, and the scaling with connection count should be much better than your original non-blocking loop.

For reference, with these optimizations, batch sends with io_uring can handle thousands of connections with minimal latency overhead, far outperforming naive loops of write() calls.

内容的提问来源于stack exchange,提问作者Johannes Ohlemacher

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 06:37:50