You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

liburing写入性能低于预期的原因排查及相关技术问题咨询

liburing写入性能低于预期的原因排查及相关技术问题咨询

我最近在做一个Linux单服务器上高速写磁盘的项目,用fio跑io_uring基准测试时能达到40GB/s以上的写入速度,但自己用liburing写的代码却只有9GB/s左右,实在头疼。先把我的情况和疑问整理出来,希望大家能帮忙分析下:

问题概述

我用下面的fio命令做基准测试,确认基于io_uring应该能达到预期的写入速度(>40GB/s):

fio --name=seqwrite --rw=write --direct=1 --ioengine=io_uring --bs=128k --numjobs=4 --size=100G --runtime=300 --directory=/mnt/md0/ --iodepth=128 --buffered=0 --numa_cpu_nodes=0 --sqthread_poll=1  --hipri=1

但用liburing写的代码却只能跑到约9GB/s,我怀疑是不是liburing有额外开销,但想先确认自己的实现逻辑有没有问题。

我的实现方式

  • 基于liburing库开发
  • 启用了提交队列轮询(SQPoll)功能
  • 没有使用writev()的分散/聚合IO,而是用普通write()提交请求(试过分散聚合,性能没明显提升)
  • 采用多线程架构,每个线程对应一个独立的io_uring实例

额外信息

  • 试过单线程版本的简化代码,性能和多线程版本差不多
  • 调试器显示线程数符合NUM_JOBS宏的设置,但不清楚内核为SQPoll创建的线程情况
  • 线程数超过2个时,性能反而下降
  • 服务器有96个CPU核心
  • 写入目标是RAID0配置的磁盘
  • 用bpftrace -e 'tracepoint:io_uring:io_uring_submit_sqe {printf("%s(%d)\n", comm, pid);}'跟踪,能看到内核的SQPoll线程在活动
  • 已验证写入磁盘的数据大小和内容完全符合预期
  • 试过创建ring时用IORING_SETUP_ATTACH_WQ标记,结果反而变慢
  • 测试过多种块大小,128k是最优选择

我的疑问

  1. 我以为每个ring实例会对应一个内核SQPoll线程,但不知道怎么验证这个机制是否真的生效?能不能默认信任它正常工作?
  2. 为什么线程数超过2个后性能会下降?是多个线程写同一个文件产生了竞争?还是其实只有一个SQPoll线程在处理所有ring的请求,导致过载?
  3. 有没有其他liburing的标记或配置选项我没用到,能提升性能?
  4. 是不是该放弃liburing,直接用io_uring的系统调用?

简化版代码示例

以下是简化后的代码,去掉了错误处理逻辑,但性能和完整版一致:

main函数

#include <fcntl.h>
#include <liburing.h>
#include <cstring>
#include <thread>
#include <vector>
#include "utilities.h"

#define NUM_JOBS 4 // number of single-ring threads
#define QUEUE_DEPTH 128 // size of each ring
#define IO_BLOCK_SIZE 128 * 1024 // write block size
#define WRITE_SIZE (IO_BLOCK_SIZE * 10000) // Total number of bytes to write
#define FILENAME  "/mnt/md0/test.txt" // File to write to

char incomingData[WRITE_SIZE]; // Will contain the data to write to disk

int main()
{
    // Initialize variables
    std::vector<std::thread> threadPool;
    std::vector<io_uring*> ringPool;
    io_uring_params params;
    int fds[2];
    int bytesPerThread = WRITE_SIZE / NUM_JOBS;
    int bytesRemaining = WRITE_SIZE % NUM_JOBS;
    int bytesAssigned = 0;

    utils::generate_data(incomingData, WRITE_SIZE); // this just fills the incomingData buffer with known data

    // Open the file, store its descriptor
    fds[0] = open(FILENAME, O_WRONLY | O_TRUNC | O_CREAT);

    // initialize Rings
    ringPool.resize(NUM_JOBS);
    for (int i = 0; i < NUM_JOBS; i++)
    {
        io_uring* ring = new io_uring;
        // Configure the io_uring parameters and init the ring
        memset(&params, 0, sizeof(params));
        params.flags |= IORING_SETUP_SQPOLL;
        params.sq_thread_idle = 2000;
        io_uring_queue_init_params(QUEUE_DEPTH, ring, &params);
        io_uring_register_files(ring, fds, 1); // required for sq polling

        // Add the ring to the pool
        ringPool.at(i) = ring;
    }

    // Spin up threads to write to the file
    threadPool.resize(NUM_JOBS);
    for (int i = 0; i < NUM_JOBS; i++)
    {
        int bytesToAssign = (i != NUM_JOBS - 1) ? bytesPerThread : bytesPerThread + bytesRemaining;
        threadPool.at(i) = std::thread(writeToFile, 0, ringPool[i], incomingData + bytesAssigned, bytesToAssign, bytesAssigned);
        bytesAssigned += bytesToAssign;
    }

    // Wait for the threads to finish
    for (int i = 0; i < NUM_JOBS; i++)
    {
        threadPool[i].join();
    }

    // Cleanup the rings
    for (int i = 0; i < NUM_JOBS; i++)
    {
        io_uring_queue_exit(ringPool[i]);
    }

    // Close the file
    close(fds[0]);

    return 0;
}

writeToFile()函数

void writeToFile(int fd, io_uring* ring, char* buffer, int size, int fileIndex)
{
    io_uring_cqe *cqe;
    io_uring_sqe *sqe;
    int bytesRemaining = size;
    int bytesToWrite;
    int bytesWritten = 0;
    int writesPending = 0;

    while (bytesRemaining || writesPending)
    {
        while(writesPending < QUEUE_DEPTH && bytesRemaining)
        {
            /* In this first inner loop,
            * Write up to QUEUE_DEPTH blocks to the submission queue
            */
            bytesToWrite = bytesRemaining > IO_BLOCK_SIZE ? IO_BLOCK_SIZE : bytesRemaining;

            sqe = io_uring_get_sqe(ring);
            if (!sqe) break; // if can't get a sqe, break out of the loop and wait for the next round

            io_uring_prep_write(sqe, fd, buffer + bytesWritten, bytesToWrite, fileIndex + bytesWritten);
            sqe->flags |= IOSQE_FIXED_FILE;

            writesPending++;
            bytesWritten += bytesToWrite;
            bytesRemaining -= bytesToWrite;

            if (bytesRemaining == 0) break;
        }

        io_uring_submit(ring);

        while(writesPending)
        {
            /* In this second inner loop,
            * Handle completions
            * Additional error handling removed for brevity
            * The functionality is the same as with errror handling in the case that nothing goes wrong
            */
            int status = io_uring_peek_cqe(ring, &cqe);
            if (status == -EAGAIN) break; // if no completions are available, break out of the loop and wait for the next round

            io_uring_cqe_seen(ring, cqe);
            writesPending--;
        }
    }
}

我的分析与建议

关于SQPoll线程的验证

你可以通过ps aux | grep io_uring-sq查看内核创建的SQPoll线程,每个开启SQPoll的ring对应一个以io_uring-sq命名的线程,线程名里还会包含进程ID和绑定的CPU核心信息。另外也可以查看/proc/<你的进程PID>/task目录,里面会列出进程的所有线程,包括这些SQPoll线程。只要正确设置了IORING_SETUP_SQPOLL标记,内核都会创建对应线程,不用过度怀疑,但验证一下更放心。

多线程性能下降的原因

  1. 文件系统与RAID并行度限制:虽然你是分片写入无重叠,但RAID0的并行度取决于底层磁盘数量,如果线程数超过磁盘数量,反而会导致IO请求排队;另外文件系统的元数据更新可能存在隐性竞争。
  2. NUMA亲和性缺失:你的服务器有96核,大概率是NUMA架构。fio里指定了--numa_cpu_nodes=0,但你的代码没做任何CPU绑定,可能导致线程或SQPoll线程跨NUMA节点访问内存,带来额外开销。
  3. SQPoll线程的CPU竞争:如果多个SQPoll线程被调度到同一个核心,会互相抢占CPU资源,降低请求处理效率。

可以尝试的优化选项

  1. 添加直接IO标记:你的fio命令用了--direct=1,但代码里打开文件时没加O_DIRECT!这是最关键的差异!直接IO可以绕过页缓存,减少内核拷贝和缓存管理开销,对大文件顺序写入性能影响极大。修改open参数为O_WRONLY | O_TRUNC | O_CREAT | O_DIRECT,同时确保缓冲区对齐(128k一般符合要求)。
  2. 设置CPU亲和性:在创建ring时,通过params.sq_thread_cpu指定SQPoll线程绑定到特定CPU核心;同时用std::thread::native_handle()结合sched_setaffinity把业务线程绑定到对应NUMA节点的核心,避免跨节点访问。
  3. 启用IORING_SETUP_IOPOLL:如果你的磁盘是NVMe这类支持IO轮询的设备,添加这个标记可以减少中断开销,进一步提升性能。
  4. 优化请求提交逻辑:尽量填满SQ队列后再调用io_uring_submit,减少系统调用次数;用io_uring_submit_and_wait替代先submit再peek_cqe的逻辑,减少用户态与内核态的切换开销。

要不要放弃liburing?

完全没必要!liburing只是io_uring系统调用的轻量封装,本身开销极小,性能和直接调用系统调用几乎无差异。你的性能问题大概率是因为配置遗漏(比如没开O_DIRECT),而非liburing本身的问题,调整好配置后,liburing代码完全能达到fio的性能水平。

备注:内容来源于stack exchange,提问作者Smitch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.23 09:22:42