使用Nsight Compute分析共享内存原子内核时出现启动失败错误
问题
我使用NVIDIA博客中的全局原子与共享原子对比代码,用Nsight Compute CLI做性能分析时,共享原子内核histogram_smem_atomics出现LaunchFailed错误,相关报错日志及main函数代码如下,请问错误原因是什么?
报错日志
==PROF== Connected to process 16078 ==PROF== Profiling "histogram_gmem_atomics" - 0: 0%....50%....100% - 1 pass ==PROF== Profiling "histogram_smem_atomics" - 1: 0%....50%....100% - 1 pass ==ERROR== LaunchFailed ==ERROR== LaunchFailed ==PROF== Trying to shutdown target application ==ERROR== The application returned an error code (9). ==ERROR== An error occurred while trying to profile. ==WARNING== Found outstanding GPU clock reset, trying to revert...Success. [16078] histogram@127.0.0.1 histogram_gmem_atomics(const IN_TYPE *, int, int, unsigned int *), 2023-Mar-09 12:55:43, Context 1, Stream 7 Section: Command line profiler metrics ---------------------------------------------------------------------- --------------- ------------------------------ dram__bytes.sum.per_second Gbyte/second 13,98 ---------------------------------------------------------------------- --------------- ------------------------------ histogram_smem_atomics(const IN_TYPE *, int, int, unsigned int *), 2023-Mar-09 12:55:43, Context 1, Stream 7 Section: Command line profiler metrics ---------------------------------------------------------------------- --------------- ------------------------------ dram__bytes.sum.per_second byte/second (!) nan ---------------------------------------------------------------------- --------------- ------------------------------
main函数代码
#define NUM_BINS 480 #define NUM_PARTS 48 struct IN_TYPE { int x; int y; int z; }; int main(){ int height = 480; int width = height; auto nThread = 16; auto nBlock = (height) / nThread; IN_TYPE* h_in_image, *d_in_image; unsigned int* d_out_image; h_in_image = (IN_TYPE *)malloc(height*width * sizeof(IN_TYPE)); cudaMalloc(&d_in_image, height*width * sizeof(IN_TYPE)); cudaMalloc(&d_out_image, height*width * sizeof(unsigned int)); for (int n = 0; n < (height*width); n++) { h_in_image[n].x = rand()%10; h_in_image[n].y = rand()%10; h_in_image[n].z = rand()%10; } cudaMemcpy(d_in_image, h_in_image, height*width * sizeof(IN_TYPE), cudaMemcpyHostToDevice); histogram_gmem_atomics<<<nBlock, nThread>>>(d_in_image, width, height, d_out_image); cudaDeviceSynchronize(); // not copying the results back as of now histogram_smem_atomics<<<nBlock, nThread>>>(d_in_image, width, height, d_out_image); cudaDeviceSynchronize(); }
错误原因分析
- 未指定共享内存分配大小:
histogram_smem_atomics内核依赖线程块级别的共享内存存储局部直方图,通常需要分配NUM_BINS * sizeof(unsigned int)大小的共享内存。但你的代码启动内核时没有添加共享内存参数,正确的启动方式应该是histogram_smem_atomics<<<nBlock, nThread, NUM_BINS * sizeof(unsigned int)>>>(...)。未分配共享内存时,内核访问共享内存会直接触发启动失败。 - 输出缓冲区尺寸不匹配:你为
d_out_image分配了和输入图像同尺寸的内存,但共享原子直方图内核的输出应该是NUM_BINS个区间的计数,只需要NUM_BINS * sizeof(unsigned int)的空间。当内核往超出分配范围的地址写数据时,会触发内存访问错误,导致内核启动失败。 - 缺少错误检查机制:代码中没有在
cudaMalloc、cudaMemcpy、内核启动后调用cudaGetLastError()、cudaPeekAtLastError()等接口检查错误,无法提前捕获共享内存未分配、内存越界这类问题,只能通过Nsight的报错间接发现。
内容的提问来源于stack exchange,提问作者yolo_ML
相关产品推荐
相关产品推荐

