关于fio测试中seqwrite与randread的blktrace块层差异问询
First, let’s recap your test setup for context:
I purchased a virtual server with 8 vCPUs, 16GB RAM, and a 500G SSD volume backed by Ceph RBD. I used the
fiotool to test IO performance, and captured block layer IO traces withblktraceduring testing to better understand results.Sequential write test command:
fio --filename=/dev/vdc --ioengine=libaio --bs=4k --rw=write --size=8G --iodepth=64 --numjobs=8 --direct=1 --runtime=960 --name=seqwrite --group_reportingRandom read test command:
fio --filename=/dev/vdc --ioengine=libaio --bs=4k --rw=randread --size=8G --iodepth=64 --numjobs=8 --direct=1 --runtime=960 --name=randread --group_reporting
Before diving into the differences, let’s clarify the key blktrace phases you’re asking about:
- I2D: The time between when an IO request is submitted to the block device driver (
Ievent) and when the driver dispatches it to the hardware queue (Devent) - Q2M: The time between when a request enters the block layer queue (
Qevent) and when the IO scheduler selects it for processing (Mevent)
Why does randread show a large number of I2D phases, but seqwrite does not?
This difference comes down to how sequential vs. random IO interacts with your SSD storage and the Ceph RBD backend:
Random read overhead:
- Each 4k random read targets a non-contiguous spot on the SSD. Even with fast SSDs, random access requires navigating to different NAND flash pages, which adds latency compared to sequential operations.
- On the Ceph side, random reads map to scattered objects across multiple OSDs. The RBD driver has to resolve each request to the correct OSD, fetch the data, and assemble it—this extra processing time makes the I2D phase clearly visible in blktrace.
- Since every random read is independent and requires this extra work, you see a large number of distinct I2D phases.
Sequential write efficiency:
- Sequential 4k writes are contiguous, so the SSD can process them in large, optimized batches (reducing write amplification and leveraging sequential NAND write speeds).
- Ceph RBD is built to handle sequential writes efficiently: it pre-allocates contiguous backend blocks, aggregates small writes into larger objects, and minimizes request resolution overhead.
- The driver can quickly batch and dispatch these contiguous requests, making the I2D phase extremely short—so short that it doesn’t register as a measurable, distinct phase in your blktrace output.
Why does randread lack a Q2M phase, while seqwrite has it?
This ties to Linux IO scheduler behavior and how your workload interacts with the queue:
IO scheduler priorities:
- Most Linux IO schedulers (like
mq-deadline, the default for many SSD setups) prioritize read requests over writes. Reads are latency-sensitive, so the scheduler will immediately pick up read requests from the queue (Q) and move them to processing (M) without waiting. There’s no delay here, so the Q2M phase doesn’t show up. - For writes, the scheduler often uses batching or sorting to maximize sequential throughput. Even though your write workload is already sequential, the scheduler may hold onto incoming requests briefly to build larger batches—this creates a measurable Q2M phase as requests wait in the queue to be grouped.
- Most Linux IO schedulers (like
Queue dynamics with your workload:
- Your
iodepth=64andnumjobs=8setup floods the queue with sequential write requests. The scheduler waits a tiny window to ensure it has a full batch of contiguous writes to dispatch, leading to visible Q2M time. - For random reads, even with the same iodepth, each request is independent and scattered. The scheduler doesn’t need to batch or sort them—each request is ready to process as soon as it enters the queue, so there’s no measurable delay between
QandMevents.
- Your
内容的提问来源于stack exchange,提问作者Ning

