OpenCL显示GPU为2单元但实际有384核?任务分配解析
Great questions! Let's break them down clearly:
The confusion here stems from two distinct hardware-level concepts: Compute Units (CUs) (what OpenCL is showing) and CUDA Cores (the count you see in CudaZ and clinfo).
- For NVIDIA GPUs, an OpenCL Compute Unit directly maps to a Streaming Multiprocessor (SM)—a higher-level hardware cluster that bundles smaller execution cores. Your GT 730 (specifically the Kepler GK208 variant) has exactly 2 SMs, which is why OpenCL logs show "2 units".
- Each of those SMs contains 192 individual CUDA Cores. 2 SMs × 192 cores = 384 total cores, which matches the number reported by CudaZ and
clinfo.
Think of it like this: Compute Units are the "teams" of cores, while CUDA Cores are the individual team members. OpenCL exposes the team count first, while tools like CudaZ focus on the total number of individual execution units.
OpenCL is built for data-parallel execution, which aligns perfectly with how GPU cores are optimized to work. Here's a straightforward breakdown:
- Work-Items: These are the smallest execution units in OpenCL, and each work-item gets assigned to a single CUDA Core. In most practical use cases, each work-item processes a unique piece of data—like one element in an array, one pixel in an image, or one sample in a dataset.
- Work-Groups: Work-items are grouped into work-groups (you define the size when launching your kernel). All work-items in a work-group run on the same Compute Unit (SM), sharing fast local memory and being able to synchronize with each other if needed.
- SIMT Execution: While each work-item handles distinct data, all work-items in a "warp" (a sub-group of 32 work-items on NVIDIA GPUs) execute the same instruction at the same time. This is called Single Instruction, Multiple Threads (SIMT)—it's efficient for GPU hardware, but the key takeaway is each core still processes unique data.
To sum it up: Different cores (work-items) handle different data, even though they might run the same kernel code. You'd only have multiple cores executing the same process on identical data if you explicitly design your kernel to do so—which is rarely useful for GPU acceleration.
内容的提问来源于stack exchange,提问作者Jaffer Wilson

