关于CUDA stream、DMA engine与Async Engine的工作机制及单Async Engine场景下传输顺序异常的问询
Hey there! Let's break down this curious behavior you're seeing on your Jetson AGX Xavier. First, let's align on what's happening:
你的预期与实际现象
You expected that with asyncEngineCount=1, async transfers should strictly follow a serial order: H2D[0] → D2H[0] → H2D[1] → D2H[1] → ... because the single DMA engine can only handle one transfer at a time. But your Nsight Systems profile shows H2D[1] runs concurrently with D2H[0], while all subsequent transfers follow the expected serial pattern.
核心原因:Jetson Xavier的硬件特性与命令调度逻辑
Let's unpack this from two key angles: hardware capabilities of the Xavier, and how the GPU command scheduler processes your submitted work.
1. Jetson Xavier的DMA引擎并非严格单任务
First, a critical detail: the asyncEngineCount field behaves differently on Jetson devices compared to desktop GPUs. For your AGX Xavier, asyncEngineCount=1 doesn't mean the DMA engine can only handle one transfer at a time. Instead, it refers to a unified DMA engine that supports bidirectional concurrent transfers (one H2D and one D2H transfer running simultaneously), as long as the transfers target independent memory regions (which yours do, thanks to your chunked array and pinned host memory from cudaMallocHost).
This is a design choice for Jetson's integrated architecture: the Unified Memory Controller (UMC) can handle simultaneous reads (D2H) and writes (H2D) to/from pinned host memory without CPU intervention.
2. Command submission order vs. execution scheduling
Your CPU submits commands in a strict sequential loop:
H2D[0] → kernel[0] → D2H[0] → H2D[1] → kernel[1] → D2H[1] → H2D[2] → kernel[2] → D2H[2] → ...
The GPU command scheduler doesn't just execute commands in submission order—it skips commands that can't run (due to stream dependencies) and prioritizes commands that are ready. Here's the play-by-play for your case:
- Step 1:
H2D[0]starts and completes, thenkernel[0]begins running on the SMs. - Step 2: You've already submitted
H2D[1](since the CPU loop runs synchronously). SinceH2D[1]is in an independent stream and the DMA engine is idle (the SMs are busy withkernel[0]), the scheduler startsH2D[1]. - Step 3: When
kernel[0]finishes,D2H[0]becomes ready. Now the DMA engine can handle bothD2H[0](D2H direction) andH2D[1](H2D direction) simultaneously—this is the bidirectional concurrency we mentioned earlier. - Step 4: For subsequent transfers (e.g.,
H2D[2]andD2H[1]),D2H[1]is submitted beforeH2D[2]in the command queue. Whenkernel[1]finishes,D2H[1]becomes ready first (it's earlier in the queue). The scheduler prioritizes it overH2D[2], soD2H[1]runs to completion beforeH2D[2]starts. This creates the serial pattern you see for the rest of the transfers.
为什么后面的传输不再并发?
Once D2H[0] and H2D[1] finish, the command queue has D2H[1] (submitted before H2D[2]) waiting for kernel[1] to complete. When kernel[1] finishes, D2H[1] is the first ready command in the queue—so the scheduler runs it immediately. H2D[2] has to wait for D2H[1] to finish because the scheduler processes commands in submission order when multiple commands become ready around the same time.
验证这个解释
If you want to test this, try reordering your command submission:
- Submit all
H2Dtransfers first, then all kernels, then allD2Htransfers. You'll see allH2Dtransfers run serially (since they're same-direction), kernels run concurrently, andD2Htransfers run serially. - Or, swap the order of
D2H[i]andH2D[i+1]in submission (e.g., submitD2H[0], thenH2D[1], thenkernel[0], etc.). You'll likely seeD2H[0]run first, thenH2D[1], since they're same-queue submission order.
总结
The key takeaway is that asyncEngineCount=1 on Jetson Xavier doesn't mean no concurrent transfers—it means no independent H2D/D2H engines, but the unified engine can handle bidirectional concurrency. The one-off concurrent pair (D2H[0] and H2D[1]) happens because H2D[1] is ready to run while D2H[0] is still waiting on its kernel, and the scheduler takes advantage of the bidirectional capability. Subsequent transfers follow submission order because the earlier command in the queue becomes ready first.
内容来源于stack exchange

