You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kepler GPU Warp调度疑问:GK110 SM中64个闲置CUDA核心的用途

Understanding GK110 SM SP Utilization: What Happens to the "Extra" 64 Cores?

Awesome question—this is a super common point of confusion when digging into the GK110’s SM architecture, especially when you’re matching up per-cycle dispatch capacity to the total number of SP cores. Let’s break down why those 64 seemingly unused cores aren’t actually sitting idle:

  • Pipeline Depth Drives Overlapped Execution
    The 192 SP cores in a GK110 SM are deeply pipelined (roughly 10-12 cycles of latency for single-precision operations). When the 4 warp schedulers dispatch 4 warps (128 threads) in one cycle, those threads’ instructions kick off at the start of the pipeline. In the next cycle, the schedulers can dispatch another 4 warps, and their instructions slot into the next stage of the pipeline. Over time, all 192 SPs get filled with instructions from different warp batches at various pipeline stages. As long as you have enough active warps (high thread occupancy) to keep the pipeline fed, every single SP is contributing to sustained throughput.

  • Dual Dispatch Supports Mixed (and Heavy) Workloads
    Each warp scheduler has dual dispatch units, meaning it can send two different types of instructions to a single warp in one cycle. For example, one dispatch might send a single-precision arithmetic task to the SPs, while the other sends a load/store or special function unit (SFU) task to dedicated hardware. But when your workload is all single-precision operations, the schedulers prioritize dispatching SP instructions across multiple warps. The extra 64 SPs come into play here by handling instructions from warps dispatched in earlier cycles, keeping the entire SM saturated with work.

  • Core Count is Optimized for Sustained Throughput
    The GK110’s SM design is all about maximizing sustained throughput, not just using every core in a single cycle. The 4 warp schedulers are built to keep as many warps active as possible, leveraging the pipeline to overlap execution. The 192 SP count is sized to handle the steady stream of instructions from multiple in-flight warp batches, not just the 4 warps dispatched in one cycle. If you only had 128 SPs, the pipeline would empty faster, leading to gaps in utilization whenever warps stall (like waiting for memory or branch resolution).

In short, those "extra" 64 cores are a key part of the architecture’s throughput optimization. They’re busy cranking through instructions from warps dispatched in earlier cycles, ensuring the SM maintains peak single-precision performance when your workload has enough threads to keep the pipeline full.

内容的提问来源于stack exchange,提问作者StrikeW

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:14:18