You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何CUDA常量内存相比全局内存无性能提升?(2D卷积场景)

3×3卷积CUDA Kernel:常量内存无性能提升的原因分析

我正在对2048×2048的图像执行3×3滤波器的2D卷积操作,分别实现了基于全局内存和常量内存存储滤波器的两个CUDA Kernel版本。在RTX 3090上测试后,发现使用常量内存没有带来性能提升。已尝试大尺寸输入图像,不确定是实现/基准测试存在问题,还是当前场景复杂度不足导致的。

使用全局内存的Kernel

__global__ void gpu_conv2d_kernel(float *d_N_ptr, float *d_F_ptr, float *d_P_ptr, int n_rows, int n_cols)
{
    // Which output element this thread works on
    int out_col = blockIdx.x*blockDim.x + threadIdx.x;
    int out_row = blockIdx.y*blockDim.y + threadIdx.y;
    
    // Check if output element is valid
    if (out_row < n_rows && out_col < n_cols) 
    {
        // Result (in thread register)
        float p_val = 0.0f;
        
        // Loop over elements of the filter array
        for (int f_row = 0; f_row < 2*FILTER_RADIUS+1; f_row++) 
        {
            for (int f_col = 0; f_col < 2*FILTER_RADIUS+1; f_col++) 
            {
                // Input element to filter element mapping
                int in_row = out_row + (f_row - FILTER_RADIUS);
                int in_col = out_col + (f_col - FILTER_RADIUS);
                        
                // Boundary check
                if (in_row >= 0 && in_row < n_rows && in_col >= 0 && in_col < n_cols) 
                    p_val += d_F_ptr[f_row*(2*FILTER_RADIUS+1) + f_col] * d_N_ptr[in_row*n_cols + in_col];
                }
        }
        d_P_ptr[out_row*n_cols + out_col] = p_val;
    }
}

使用常量内存的Kernel

#define FILTER_RADIUS 1
extern __constant__ float d_F[(2*FILTER_RADIUS+1)*(2*FILTER_RADIUS+1)];

__global__ void gpu_conv2d_constMem_kernel(float *d_N_ptr, float *d_P_ptr, int n_rows, int n_cols)
{
    // Which output element this thread works on
    int out_col = blockIdx.x*blockDim.x + threadIdx.x;
    int out_row = blockIdx.y*blockDim.y + threadIdx.y;
    
    // Check if output element is valid
    if (out_row < n_rows && out_col < n_cols) 
    {
        // Result (in thread register)
        float p_val = 0.0f;
        
        // Loop over elements of the filter array
        for (int f_row = 0; f_row < 2*FILTER_RADIUS+1; f_row++) 
        {
            for (int f_col = 0; f_col < 2*FILTER_RADIUS+1; f_col++) 
            {
                // Input element to filter element mapping
                int in_row = out_row + (f_row - FILTER_RADIUS);
                int in_col = out_col + (f_col - FILTER_RADIUS);
                
                // Boundary check
                if (in_row >= 0 && in_row < n_rows && in_col >= 0 && in_col < n_cols) 
                    p_val += d_F[f_row*(2*FILTER_RADIUS+1)+f_col] * d_N_ptr[in_row*n_cols + in_col];
            }
        }
        d_P_ptr[out_row*n_cols + out_col] = p_val;
    }
}

可能的原因

  • 全局内存缓存抵消优势:RTX 3090的L1/L2缓存对全局内存访问优化出色。3×3滤波器仅9个元素,多线程访问时会被快速缓存,后续访问命中缓存的延迟和常量内存几乎一致,常量内存的低延迟优势无法体现。
  • 卷积瓶颈在输入图像读取:2048×2048图像的每个输出像素需读取9个输入像素(边界除外),输入图像的全局内存访问才是性能瓶颈。滤波器读取量占比极小,即使从全局内存读取也不会影响整体性能。
  • 常量内存特性未充分利用:常量内存适合大量线程高频读取相同数据的场景,但3×3滤波器元素少,线程读取滤波器的次数远少于输入图像,其高带宽特性无法发挥。
  • 基准测试误差:若测试未排除内存分配、数据传输的干扰,或测试次数不足,性能差异可能被噪声掩盖。需确保单独测试Kernel执行时间,并多次取平均。

内容的提问来源于stack exchange,提问作者Tushar Gautam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 23:15:55