关于cuBLAS批处理GEMM性能异常的技术咨询
Great question—this counterintuitive performance jump is actually a result of how CuBLAS adapts its execution strategy to the problem size, combined with how the GPU's streaming multiprocessors (SMs) utilize parallel workloads. Let's break down what's happening here, using your GTX 1080Ti test results as context.
Key Reasons for the Performance Jump
1. CuBLAS's Adaptive Kernel Selection
CuBLAS doesn't rely on a single algorithm for all matrix multiplication tasks. It uses a heuristic-driven scheduler that picks the optimal kernel implementation based on factors like individual matrix dimensions, total batch count, and data layout.
For small batch sizes (100, 1000, 10000), CuBLAS likely uses kernels optimized for small, individual matrix multiplications. These kernels don't fully saturate all 28 SMs on your GTX 1080Ti—each SM only gets a tiny workload, leaving most of the GPU's compute resources idle.
Once the batch size hits a critical threshold (100000 in your tests), CuBLAS switches to kernels designed for massively batched workloads. These kernels distribute the batch across all SMs evenly, ensuring every compute unit is continuously busy. This shift eliminates idle time and leverages the GPU's parallel nature to its full potential.
2. GPU SM Utilization & Workload Granularity
GPUs thrive on large, parallel workloads. Let's do a quick back-of-the-envelope calculation for your 10x10 matrix test:
- For batch=10000: 10000 matrices / 28 SMs ≈ 357 matrices per SM. Each 10x10 matmul is a tiny task—SMs finish their small workloads quickly and wait for new tasks, wasting cycles on scheduling overhead.
- For batch=100000: 100000 matrices /28 SMs ≈ 3571 matrices per SM. This larger workload keeps each SM occupied continuously, minimizing scheduling overhead and maximizing compute utilization.
This is why the performance jump is even more dramatic for 10x10 matrices—their individual compute footprint is smaller, so small batches waste more GPU resources compared to large batches.
3. Memory Access Efficiency
Large batch sizes also improve memory utilization. In strided batched mode, larger batches mean more consecutive memory accesses to your d_A, d_B, and d_C buffers. This better aligns with the GPU's memory hierarchy (L1/L2 caches, global memory bandwidth), reducing cache misses and increasing the rate at which data is fed to compute units. Small batches, by contrast, have more scattered memory access patterns that don't leverage caching effectively.
Critical Bug in Your Test Code
Before diving deeper, I noticed a mistake in your memory copy directions:
cudaMemcpy(h_A,d_A,sizeof(float) * dim * dim * batch_count,cudaMemcpyDeviceToHost); cudaMemcpy(h_B,d_B,sizeof(float) * dim * dim * batch_count,cudaMemcpyDeviceToHost); cudaMemcpy(h_C,d_C,sizeof(float) * dim * dim * batch_count,cudaMemcpyDeviceToHost);
You're copying from device to host here, but you need to copy your host matrices (h_A, h_B) to the device (d_A, d_B) using cudaMemcpyHostToDevice. This bug means your device matrices are filled with uninitialized garbage data—while it doesn't affect the timing measurement (the kernel still executes the same number of operations), it makes the computed h_C values meaningless. Fix this for valid correctness testing.
How to Verify This
To confirm these hypotheses, try these steps:
- Profile with NVIDIA tools: Use
nvprofornsys profileto inspect kernel execution. Run:
You'll see different kernel names for small vs large batches, confirming CuBLAS is switching implementations. You can also check SM utilization metrics to see how much of the GPU is being used.nvprof --print-gpu-trace ./cublas.out - Test intermediate batch sizes: Try batch sizes like 50000, 150000 to find the exact threshold where the performance jump occurs. This will align with when CuBLAS switches its kernel strategy.
- Check linear scalability post-threshold: For batches larger than 100000 (like 200000, 500000), you should see linear scaling again—once all SMs are fully utilized, doubling the batch size should roughly double the execution time.
Summary
The "反常" performance boost isn't a bug—it's CuBLAS doing its job by adapting to the workload size. Small batches don't fully utilize the GPU's parallel resources, but once you hit a critical batch size, CuBLAS switches to a high-throughput execution mode that saturates all SMs and optimizes memory access. This effect is more pronounced for smaller matrices because their individual workloads are too tiny to keep the GPU busy at small batch sizes.
内容的提问来源于stack exchange,提问作者qtqt

