为何CUDA常量内存相比全局内存无性能提升?(2D卷积场景)
3×3卷积CUDA Kernel:常量内存无性能提升的原因分析
我正在对2048×2048的图像执行3×3滤波器的2D卷积操作,分别实现了基于全局内存和常量内存存储滤波器的两个CUDA Kernel版本。在RTX 3090上测试后,发现使用常量内存没有带来性能提升。已尝试大尺寸输入图像,不确定是实现/基准测试存在问题,还是当前场景复杂度不足导致的。
使用全局内存的Kernel
__global__ void gpu_conv2d_kernel(float *d_N_ptr, float *d_F_ptr, float *d_P_ptr, int n_rows, int n_cols) { // Which output element this thread works on int out_col = blockIdx.x*blockDim.x + threadIdx.x; int out_row = blockIdx.y*blockDim.y + threadIdx.y; // Check if output element is valid if (out_row < n_rows && out_col < n_cols) { // Result (in thread register) float p_val = 0.0f; // Loop over elements of the filter array for (int f_row = 0; f_row < 2*FILTER_RADIUS+1; f_row++) { for (int f_col = 0; f_col < 2*FILTER_RADIUS+1; f_col++) { // Input element to filter element mapping int in_row = out_row + (f_row - FILTER_RADIUS); int in_col = out_col + (f_col - FILTER_RADIUS); // Boundary check if (in_row >= 0 && in_row < n_rows && in_col >= 0 && in_col < n_cols) p_val += d_F_ptr[f_row*(2*FILTER_RADIUS+1) + f_col] * d_N_ptr[in_row*n_cols + in_col]; } } d_P_ptr[out_row*n_cols + out_col] = p_val; } }
使用常量内存的Kernel
#define FILTER_RADIUS 1 extern __constant__ float d_F[(2*FILTER_RADIUS+1)*(2*FILTER_RADIUS+1)]; __global__ void gpu_conv2d_constMem_kernel(float *d_N_ptr, float *d_P_ptr, int n_rows, int n_cols) { // Which output element this thread works on int out_col = blockIdx.x*blockDim.x + threadIdx.x; int out_row = blockIdx.y*blockDim.y + threadIdx.y; // Check if output element is valid if (out_row < n_rows && out_col < n_cols) { // Result (in thread register) float p_val = 0.0f; // Loop over elements of the filter array for (int f_row = 0; f_row < 2*FILTER_RADIUS+1; f_row++) { for (int f_col = 0; f_col < 2*FILTER_RADIUS+1; f_col++) { // Input element to filter element mapping int in_row = out_row + (f_row - FILTER_RADIUS); int in_col = out_col + (f_col - FILTER_RADIUS); // Boundary check if (in_row >= 0 && in_row < n_rows && in_col >= 0 && in_col < n_cols) p_val += d_F[f_row*(2*FILTER_RADIUS+1)+f_col] * d_N_ptr[in_row*n_cols + in_col]; } } d_P_ptr[out_row*n_cols + out_col] = p_val; } }
可能的原因
- 全局内存缓存抵消优势:RTX 3090的L1/L2缓存对全局内存访问优化出色。3×3滤波器仅9个元素,多线程访问时会被快速缓存,后续访问命中缓存的延迟和常量内存几乎一致,常量内存的低延迟优势无法体现。
- 卷积瓶颈在输入图像读取:2048×2048图像的每个输出像素需读取9个输入像素(边界除外),输入图像的全局内存访问才是性能瓶颈。滤波器读取量占比极小,即使从全局内存读取也不会影响整体性能。
- 常量内存特性未充分利用:常量内存适合大量线程高频读取相同数据的场景,但3×3滤波器元素少,线程读取滤波器的次数远少于输入图像,其高带宽特性无法发挥。
- 基准测试误差:若测试未排除内存分配、数据传输的干扰,或测试次数不足,性能差异可能被噪声掩盖。需确保单独测试Kernel执行时间,并多次取平均。
内容的提问来源于stack exchange,提问作者Tushar Gautam
相关产品推荐
相关产品推荐

