CUDA中cudaGetSymbolAddress使用疑问:极简代码触发段错误
嘿,我帮你揪出代码里的两个关键问题啦,正是它们导致了段错误!咱们一步步拆解修复:
问题分析与修复方案
1. cudaGetSymbolAddress 参数传错了
你调用这个函数时,第二个参数传的是 device_int——这是设备变量的值,但CUDA需要的是这个设备符号的地址,也就是 &device_int,这样它才能定位到设备上这个变量的位置。
错误写法:
cudaGetSymbolAddress((void **)&host_pointer_to_device_int, device_int);
修复后:
cudaGetSymbolAddress((void **)&host_pointer_to_device_int, &device_int);
2. 主机指针未初始化,访问了非法内存
你的 host_pointer_to_host_int 是个野指针,根本没指向任何有效的主机内存空间。直接用 cudaMemcpy 往这个指针写数据,系统自然会触发段错误来阻止非法内存访问。
这里有两种简单的修复方式:
方式一:用栈上的普通变量(推荐,更简单)
直接在主机栈上分配一个int变量,把设备变量的值复制到这个变量里:
int host_int; // 栈上分配有效内存 cudaMemcpy(&host_int, host_pointer_to_device_int, sizeof(int), cudaMemcpyDeviceToHost); cout << host_int << endl;
方式二:动态分配主机内存
如果需要用指针的话,记得先给指针分配有效的内存空间:
int *host_pointer_to_host_int = (int*)malloc(sizeof(int)); assert(host_pointer_to_host_int != nullptr); // 确保分配成功 cudaMemcpy(host_pointer_to_host_int, host_pointer_to_device_int, sizeof(int), cudaMemcpyDeviceToHost); cout << *host_pointer_to_host_int << endl; free(host_pointer_to_host_int); // 用完记得释放内存
完整修复后的代码
#include <iostream> #include <cassert> #include <cstdlib> using namespace std; __device__ int device_int; __global__ void kernel() { device_int = 1000; } int main() { kernel<<<1, 1>>>(); // 建议加上同步,确保内核执行完成再后续操作 cudaDeviceSynchronize(); assert(cudaGetLastError() == cudaSuccess); int *host_pointer_to_device_int; // 修复符号地址参数 cudaGetSymbolAddress((void **)&host_pointer_to_device_int, &device_int); assert(cudaGetLastError() == cudaSuccess); // 方式一:栈变量实现 int host_int; cudaMemcpy(&host_int, host_pointer_to_device_int, sizeof(int), cudaMemcpyDeviceToHost); assert(cudaGetLastError() == cudaSuccess); cout << host_int << endl; // 方式二:动态内存实现(注释掉方式一可以启用这个) /* int *host_pointer_to_host_int = (int*)malloc(sizeof(int)); assert(host_pointer_to_host_int != nullptr); cudaMemcpy(host_pointer_to_host_int, host_pointer_to_device_int, sizeof(int), cudaMemcpyDeviceToHost); assert(cudaGetLastError() == cudaSuccess); cout << *host_pointer_to_host_int << endl; free(host_pointer_to_host_int); */ return 0; }
额外提一句:调用内核后最好加上 cudaDeviceSynchronize(),虽然你的代码里用 cudaGetLastError() 检查了内核启动错误,但CUDA内核是异步执行的,不加同步的话,后续的内存复制操作可能在内核还没写完变量时就开始了,容易出现不确定的问题。养成同步的习惯会更稳妥~
内容的提问来源于stack exchange,提问作者Kolay.Ne
相关产品推荐
相关产品推荐

