CUDA剩余显存充足却报‘out of memory’错误求助
问题分析与解决:CUDA设备函数中cudaMalloc引发的"内存不足"错误
核心问题
你在__device__修饰的设备函数中调用了cudaMalloc,这是完全错误的用法——cudaMalloc是CUDA运行时API,仅能在主机端(CPU代码)调用,设备端(GPU内核/设备函数)无权执行该接口。你看到的"out of memory"并非真的显存不足,而是设备端调用不支持的API导致的错误反馈。
错误原因拆解
- 每个线程都会执行
generate_spike_train中的cudaMalloc,但设备环境根本不支持这个运行时API,触发的错误被错误检查函数捕获,表现为重复的"out of memory"提示。 - 即便你用了
-rdc=true编译选项(支持可重定位设备代码),也仅解决了设备函数的链接问题,无法让主机端的运行时API在设备端生效。
修复方案
1. 替换设备端内存分配接口
将设备函数中的cudaMalloc/cudaFree替换为CUDA设备端专用的malloc/free:
__device__ int* generate_spike_train() { const int n = 70000; int* spike_train = (int*)malloc(n * sizeof(int)); // 设备端malloc if (spike_train == nullptr) { printf("Device malloc failed at line %d\n", __LINE__); return nullptr; } // Do stuff; return spike_train; } __global__ void run_simulations() { int* spike_train = generate_spike_train(); if (spike_train == nullptr) return; // Do stuff. free(spike_train); // 设备端free }
2. 调整设备堆大小(可选)
设备端malloc从设备堆分配内存,默认堆大小有限。如果需要分配较大内存,可在主机端启动内核前调整堆大小:
int main(int argc, char *argv[]) { const float to_MB = 1024.0f * 1024.0f; const int n_blocks = 30; const int n_threads = 2; size_t free, total; // 调整设备堆大小为50MB(根据实际需求修改) size_t heap_size = 50 * 1024 * 1024; cudaDeviceSetLimit(cudaLimitMallocHeapSize, heap_size); cudaMemGetInfo(&free, &total); printf("Memory left: %f / %f MB\n", free / to_MB, total / to_MB); run_simulations<<<n_blocks, n_threads>>>(); cudaDeviceSynchronize(); cudaMemGetInfo(&free, &total); printf("Memory left: %f / %f MB\n", free / to_MB, total / to_MB); }
3. 修正错误检查逻辑
原__device__的check函数中调用了主机端的cudaGetErrorString,这在设备端无法执行。设备端只需检查malloc的返回值是否为nullptr即可判断分配是否成功。
验证
修改后用原命令编译运行:
$ nvcc -rdc=true test.cu $ ./a.out
此时不会再出现"out of memory"错误,显存占用也会符合预期。
内容的提问来源于stack exchange,提问作者Drade
相关产品推荐
相关产品推荐

