调用Nvidia NPPi函数后目标指针损坏,cudaMemcpy非法访问求助
NPPi卷积后cudaMemcpy非法内存访问问题排查与解决
核心原因分析
1. 卷积核未拷贝至设备内存
nppiFilterBorder_32f_C1R是GPU端执行的函数,其pKernel参数要求传入设备内存指针。你当前的kernel数组位于主机内存,GPU无法直接访问主机内存,会导致隐式错误——虽然函数返回NPP_SUCCESS,但实际GPU执行时发生非法内存访问,破坏了deviceOut的内存结构。
2. 锚点设置不符合卷积逻辑
3x3卷积核的锚点通常应设为核的中心位置(1,1)(索引从0开始),你当前设置的(0,0)会导致卷积计算时的位置偏移,虽不一定直接触发内存错误,但可能导致输出结果不符合预期,甚至在某些边界场景下间接引发内存越界。
3. 变量一致性风险
代码中hostOut = malloc(...)但后续cudaMemcpy使用的是hostOutputData,需确保hostOutputData已正确分配足够内存(大小为width * height * sizeof(float)),否则会导致主机端内存访问错误。
修复后的代码示例
#define KERNEL_WIDTH 3 #define KERNEL_HEIGHT 3 #define KERNEL_SIZE (KERNEL_WIDTH * KERNEL_HEIGHT) #define KERNEL {0.0f, 0.25f, 0.0f, 0.25f, 0.0f, 0.25f, 0.0f, 0.25f, 0.0f} float kernel[KERNEL_SIZE] = KERNEL; float *deviceKernel; // 新增设备端核指针 NppiSize kernelSize = {.width=KERNEL_WIDTH, .height=KERNEL_HEIGHT}; int width; int height; float *hostIn; // 已填充数据 float *hostOutputData; // 确保此指针已正确分配内存 float *deviceIn; float *deviceOut; // 初始化主机输出内存(如果hostOutputData未分配) hostOutputData = malloc(width * height * sizeof(float)); cudaMalloc((void **)&deviceIn, width * height * sizeof(float)); cudaMalloc((void **)&deviceOut, width * height * sizeof(float)); cudaMalloc((void **)&deviceKernel, KERNEL_SIZE * sizeof(float)); // 分配设备核内存 // 拷贝输入数据和核数据到设备 cudaMemcpy(deviceIn, hostIn, width * height * sizeof(float), cudaMemcpyHostToDevice); cudaMemcpy(deviceKernel, kernel, KERNEL_SIZE * sizeof(float), cudaMemcpyHostToDevice); // 验证初始拷贝(可选) cudaMemcpy(hostOutputData, deviceOut, width * height * sizeof(float), cudaMemcpyDeviceToHost); NppiSize inputSize = {.width=width, .height=height}; NppiSize oSizeROI = {.width=width, .height=height}; NppiPoint oAnchor = {.x=1, .y=1}; // 修改为3x3核的中心锚点 NppiPoint oSrcOffset = {.x=0, .y=0}; // 使用设备端核指针调用NPP函数 NppStatus err = nppiFilterBorder_32f_C1R(deviceIn, width * sizeof(float), inputSize, oSrcOffset, deviceOut, width*sizeof(float), oSizeROI, deviceKernel, kernelSize, oAnchor, NPP_BORDER_REPLICATE); assert(err == NPP_SUCCESS); // 检查CUDA错误(在拷贝前添加,排查中间错误) cudaError_t cudaErr = cudaGetLastError(); if (cudaErr != cudaSuccess) { printf("CUDA error before memcpy: %s\n", cudaGetErrorString(cudaErr)); } // 执行拷贝 cudaMemcpy(hostOutputData, deviceOut, width * height * sizeof(float), cudaMemcpyDeviceToHost); // 释放内存(不要遗漏) free(hostOutputData); cudaFree(deviceIn); cudaFree(deviceOut); cudaFree(deviceKernel);
额外调试建议
- 在调用NPP函数后、拷贝前,添加
cudaGetLastError()检查CUDA runtime错误——NPP的NPP_SUCCESS仅表示函数参数校验通过,不代表GPU执行无错误。 - 使用
cuda-memcheck工具运行程序,定位具体的内存访问越界位置,帮助排查深层问题。
内容的提问来源于stack exchange,提问作者Vlad Zhdanov
相关产品推荐
相关产品推荐

