You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Nsight Compute分析共享内存原子内核时出现启动失败错误

问题

我使用NVIDIA博客中的全局原子与共享原子对比代码,用Nsight Compute CLI做性能分析时,共享原子内核histogram_smem_atomics出现LaunchFailed错误,相关报错日志及main函数代码如下,请问错误原因是什么?

报错日志

==PROF== Connected to process 16078
==PROF== Profiling "histogram_gmem_atomics" - 0: 0%....50%....100% - 1 pass
==PROF== Profiling "histogram_smem_atomics" - 1: 0%....50%....100% - 1 pass

==ERROR== LaunchFailed

==ERROR== LaunchFailed
==PROF== Trying to shutdown target application
==ERROR== The application returned an error code (9).
==ERROR== An error occurred while trying to profile.
==WARNING== Found outstanding GPU clock reset, trying to revert...Success.
[16078] histogram@127.0.0.1
  histogram_gmem_atomics(const IN_TYPE *, int, int, unsigned int *), 2023-Mar-09 12:55:43, Context 1, Stream 7
    Section: Command line profiler metrics
    ---------------------------------------------------------------------- --------------- ------------------------------
    dram__bytes.sum.per_second                                                Gbyte/second                          13,98
    ---------------------------------------------------------------------- --------------- ------------------------------

  histogram_smem_atomics(const IN_TYPE *, int, int, unsigned int *), 2023-Mar-09 12:55:43, Context 1, Stream 7
    Section: Command line profiler metrics
    ---------------------------------------------------------------------- --------------- ------------------------------
    dram__bytes.sum.per_second                                                 byte/second                        (!) nan
    ---------------------------------------------------------------------- --------------- ------------------------------

main函数代码

#define NUM_BINS 480
#define NUM_PARTS 48

struct IN_TYPE
{
    int x;
    int y;
    int z;
};

int main(){
    int height = 480;
    int width = height;

    auto nThread = 16;
    auto nBlock = (height) / nThread;

    IN_TYPE* h_in_image, *d_in_image;
    unsigned int* d_out_image;
    h_in_image = (IN_TYPE *)malloc(height*width * sizeof(IN_TYPE));
    cudaMalloc(&d_in_image, height*width * sizeof(IN_TYPE));
    cudaMalloc(&d_out_image, height*width * sizeof(unsigned int));

    for (int n = 0; n < (height*width); n++)
    {
        h_in_image[n].x = rand()%10;
        h_in_image[n].y = rand()%10;
        h_in_image[n].z = rand()%10;
    }
    cudaMemcpy(d_in_image, h_in_image, height*width * sizeof(IN_TYPE), cudaMemcpyHostToDevice);

    histogram_gmem_atomics<<<nBlock, nThread>>>(d_in_image, width, height, d_out_image);
    cudaDeviceSynchronize();

// not copying the results back as of now

    histogram_smem_atomics<<<nBlock, nThread>>>(d_in_image, width, height, d_out_image);
    cudaDeviceSynchronize();

}
错误原因分析
  • 未指定共享内存分配大小:histogram_smem_atomics内核依赖线程块级别的共享内存存储局部直方图,通常需要分配NUM_BINS * sizeof(unsigned int)大小的共享内存。但你的代码启动内核时没有添加共享内存参数,正确的启动方式应该是histogram_smem_atomics<<<nBlock, nThread, NUM_BINS * sizeof(unsigned int)>>>(...)。未分配共享内存时,内核访问共享内存会直接触发启动失败。
  • 输出缓冲区尺寸不匹配:你为d_out_image分配了和输入图像同尺寸的内存,但共享原子直方图内核的输出应该是NUM_BINS个区间的计数,只需要NUM_BINS * sizeof(unsigned int)的空间。当内核往超出分配范围的地址写数据时,会触发内存访问错误,导致内核启动失败。
  • 缺少错误检查机制:代码中没有在cudaMalloc、cudaMemcpy、内核启动后调用cudaGetLastError()、cudaPeekAtLastError()等接口检查错误,无法提前捕获共享内存未分配、内存越界这类问题,只能通过Nsight的报错间接发现。

内容的提问来源于stack exchange,提问作者yolo_ML

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 16:32:49