liburing写入性能低于预期的原因排查及相关技术问题咨询
liburing写入性能低于预期的原因排查及相关技术问题咨询
我最近在做一个Linux单服务器上高速写磁盘的项目,用fio跑io_uring基准测试时能达到40GB/s以上的写入速度,但自己用liburing写的代码却只有9GB/s左右,实在头疼。先把我的情况和疑问整理出来,希望大家能帮忙分析下:
问题概述
我用下面的fio命令做基准测试,确认基于io_uring应该能达到预期的写入速度(>40GB/s):
fio --name=seqwrite --rw=write --direct=1 --ioengine=io_uring --bs=128k --numjobs=4 --size=100G --runtime=300 --directory=/mnt/md0/ --iodepth=128 --buffered=0 --numa_cpu_nodes=0 --sqthread_poll=1 --hipri=1
但用liburing写的代码却只能跑到约9GB/s,我怀疑是不是liburing有额外开销,但想先确认自己的实现逻辑有没有问题。
我的实现方式
- 基于liburing库开发
- 启用了提交队列轮询(SQPoll)功能
- 没有使用
writev()的分散/聚合IO,而是用普通write()提交请求(试过分散聚合,性能没明显提升) - 采用多线程架构,每个线程对应一个独立的io_uring实例
额外信息
- 试过单线程版本的简化代码,性能和多线程版本差不多
- 调试器显示线程数符合
NUM_JOBS宏的设置,但不清楚内核为SQPoll创建的线程情况 - 线程数超过2个时,性能反而下降
- 服务器有96个CPU核心
- 写入目标是RAID0配置的磁盘
- 用
bpftrace -e 'tracepoint:io_uring:io_uring_submit_sqe {printf("%s(%d)\n", comm, pid);}'跟踪,能看到内核的SQPoll线程在活动 - 已验证写入磁盘的数据大小和内容完全符合预期
- 试过创建ring时用
IORING_SETUP_ATTACH_WQ标记,结果反而变慢 - 测试过多种块大小,128k是最优选择
我的疑问
- 我以为每个ring实例会对应一个内核SQPoll线程,但不知道怎么验证这个机制是否真的生效?能不能默认信任它正常工作?
- 为什么线程数超过2个后性能会下降?是多个线程写同一个文件产生了竞争?还是其实只有一个SQPoll线程在处理所有ring的请求,导致过载?
- 有没有其他liburing的标记或配置选项我没用到,能提升性能?
- 是不是该放弃liburing,直接用io_uring的系统调用?
简化版代码示例
以下是简化后的代码,去掉了错误处理逻辑,但性能和完整版一致:
main函数
#include <fcntl.h> #include <liburing.h> #include <cstring> #include <thread> #include <vector> #include "utilities.h" #define NUM_JOBS 4 // number of single-ring threads #define QUEUE_DEPTH 128 // size of each ring #define IO_BLOCK_SIZE 128 * 1024 // write block size #define WRITE_SIZE (IO_BLOCK_SIZE * 10000) // Total number of bytes to write #define FILENAME "/mnt/md0/test.txt" // File to write to char incomingData[WRITE_SIZE]; // Will contain the data to write to disk int main() { // Initialize variables std::vector<std::thread> threadPool; std::vector<io_uring*> ringPool; io_uring_params params; int fds[2]; int bytesPerThread = WRITE_SIZE / NUM_JOBS; int bytesRemaining = WRITE_SIZE % NUM_JOBS; int bytesAssigned = 0; utils::generate_data(incomingData, WRITE_SIZE); // this just fills the incomingData buffer with known data // Open the file, store its descriptor fds[0] = open(FILENAME, O_WRONLY | O_TRUNC | O_CREAT); // initialize Rings ringPool.resize(NUM_JOBS); for (int i = 0; i < NUM_JOBS; i++) { io_uring* ring = new io_uring; // Configure the io_uring parameters and init the ring memset(¶ms, 0, sizeof(params)); params.flags |= IORING_SETUP_SQPOLL; params.sq_thread_idle = 2000; io_uring_queue_init_params(QUEUE_DEPTH, ring, ¶ms); io_uring_register_files(ring, fds, 1); // required for sq polling // Add the ring to the pool ringPool.at(i) = ring; } // Spin up threads to write to the file threadPool.resize(NUM_JOBS); for (int i = 0; i < NUM_JOBS; i++) { int bytesToAssign = (i != NUM_JOBS - 1) ? bytesPerThread : bytesPerThread + bytesRemaining; threadPool.at(i) = std::thread(writeToFile, 0, ringPool[i], incomingData + bytesAssigned, bytesToAssign, bytesAssigned); bytesAssigned += bytesToAssign; } // Wait for the threads to finish for (int i = 0; i < NUM_JOBS; i++) { threadPool[i].join(); } // Cleanup the rings for (int i = 0; i < NUM_JOBS; i++) { io_uring_queue_exit(ringPool[i]); } // Close the file close(fds[0]); return 0; }
writeToFile()函数
void writeToFile(int fd, io_uring* ring, char* buffer, int size, int fileIndex) { io_uring_cqe *cqe; io_uring_sqe *sqe; int bytesRemaining = size; int bytesToWrite; int bytesWritten = 0; int writesPending = 0; while (bytesRemaining || writesPending) { while(writesPending < QUEUE_DEPTH && bytesRemaining) { /* In this first inner loop, * Write up to QUEUE_DEPTH blocks to the submission queue */ bytesToWrite = bytesRemaining > IO_BLOCK_SIZE ? IO_BLOCK_SIZE : bytesRemaining; sqe = io_uring_get_sqe(ring); if (!sqe) break; // if can't get a sqe, break out of the loop and wait for the next round io_uring_prep_write(sqe, fd, buffer + bytesWritten, bytesToWrite, fileIndex + bytesWritten); sqe->flags |= IOSQE_FIXED_FILE; writesPending++; bytesWritten += bytesToWrite; bytesRemaining -= bytesToWrite; if (bytesRemaining == 0) break; } io_uring_submit(ring); while(writesPending) { /* In this second inner loop, * Handle completions * Additional error handling removed for brevity * The functionality is the same as with errror handling in the case that nothing goes wrong */ int status = io_uring_peek_cqe(ring, &cqe); if (status == -EAGAIN) break; // if no completions are available, break out of the loop and wait for the next round io_uring_cqe_seen(ring, cqe); writesPending--; } } }
我的分析与建议
关于SQPoll线程的验证
你可以通过ps aux | grep io_uring-sq查看内核创建的SQPoll线程,每个开启SQPoll的ring对应一个以io_uring-sq命名的线程,线程名里还会包含进程ID和绑定的CPU核心信息。另外也可以查看/proc/<你的进程PID>/task目录,里面会列出进程的所有线程,包括这些SQPoll线程。只要正确设置了IORING_SETUP_SQPOLL标记,内核都会创建对应线程,不用过度怀疑,但验证一下更放心。
多线程性能下降的原因
- 文件系统与RAID并行度限制:虽然你是分片写入无重叠,但RAID0的并行度取决于底层磁盘数量,如果线程数超过磁盘数量,反而会导致IO请求排队;另外文件系统的元数据更新可能存在隐性竞争。
- NUMA亲和性缺失:你的服务器有96核,大概率是NUMA架构。fio里指定了
--numa_cpu_nodes=0,但你的代码没做任何CPU绑定,可能导致线程或SQPoll线程跨NUMA节点访问内存,带来额外开销。 - SQPoll线程的CPU竞争:如果多个SQPoll线程被调度到同一个核心,会互相抢占CPU资源,降低请求处理效率。
可以尝试的优化选项
- 添加直接IO标记:你的fio命令用了
--direct=1,但代码里打开文件时没加O_DIRECT!这是最关键的差异!直接IO可以绕过页缓存,减少内核拷贝和缓存管理开销,对大文件顺序写入性能影响极大。修改open参数为O_WRONLY | O_TRUNC | O_CREAT | O_DIRECT,同时确保缓冲区对齐(128k一般符合要求)。 - 设置CPU亲和性:在创建ring时,通过
params.sq_thread_cpu指定SQPoll线程绑定到特定CPU核心;同时用std::thread::native_handle()结合sched_setaffinity把业务线程绑定到对应NUMA节点的核心,避免跨节点访问。 - 启用IORING_SETUP_IOPOLL:如果你的磁盘是NVMe这类支持IO轮询的设备,添加这个标记可以减少中断开销,进一步提升性能。
- 优化请求提交逻辑:尽量填满SQ队列后再调用
io_uring_submit,减少系统调用次数;用io_uring_submit_and_wait替代先submit再peek_cqe的逻辑,减少用户态与内核态的切换开销。
要不要放弃liburing?
完全没必要!liburing只是io_uring系统调用的轻量封装,本身开销极小,性能和直接调用系统调用几乎无差异。你的性能问题大概率是因为配置遗漏(比如没开O_DIRECT),而非liburing本身的问题,调整好配置后,liburing代码完全能达到fio的性能水平。
备注:内容来源于stack exchange,提问作者Smitch
相关产品推荐
相关产品推荐

