Ubuntu24.04下cudaGetDevice调用失败,提示out of memory求助
CUDA程序调用cudaGetDevice返回out of memory问题排查与解决
问题描述
我编写了如下CUDA程序:
#include <cuda_runtime_api.h> #include <iostream> // 原代码缺少该头文件,需补充才能正常使用std::cout int main() { int id = -1; cudaError_t state = cudaGetDevice(&id); // 建议用cudaError_t替代int存储错误状态 if (state != cudaSuccess) { std::cout << "Failed!" << std::endl; std::cout << cudaGetErrorString(state) << std::endl; } return 0; }
运行后输出Failed!,添加错误信息打印后,得到具体错误为out of memory。此前尝试排查CUDA驱动与编译器版本不匹配的问题,但无效果。
环境信息
nvcc版本:
nvcc: NVIDIA (R) Cuda compiler driver Copyright (c) 2005-2024 NVIDIA Corporation Built on Tue_Oct_29_23:50:19_PDT_2024 Cuda compilation tools, release 12.6, V12.6.85 Build cuda_12.6.r12.6/compiler.35059454_0
NVIDIA驱动与GPU状态:
⋊> /u/l/c/bin nvidia-smi (base) 23:20:24 Mon Dec 2 23:20:47 2024 +-----------------------------------------------------------------------------------------+ | NVIDIA-SMI 560.35.03 Driver Version: 560.35.03 CUDA Version: 12.6 | |-----------------------------------------+------------------------+----------------------+ | GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+========================+======================| | 0 NVIDIA GeForce RTX 3060 Ti Off | 00000000:01:00.0 On | N/A | | 0% 41C P8 16W / 225W | 83MiB / 8192MiB | 41% Default | | | | N/A | +-----------------------------------------+------------------------+----------------------+ +-----------------------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=========================================================================================| | 0 N/A N/A 2908 G /usr/lib/xorg/Xorg 69MiB | +-----------------------------------------------------------------------------------------+
解决方案
- 排查隐藏内存占用:用
nvidia-smi pmon查看实时GPU进程内存使用,或用ps aux | grep -E "(nvidia|Xorg)"找出所有关联进程,关闭不必要的GPU占用程序(如后台渲染、闲置AI推理进程等)。 - 调整CUDA初始化参数:在调用
cudaGetDevice前添加以下代码,限制初始化阶段的内存分配:cudaSetDeviceFlags(cudaDeviceScheduleBlockingSync); cudaDeviceSetLimit(cudaLimitMallocHeapSize, 1 * 1024 * 1024); // 将堆内存限制为1MB - 减少GUI环境占用:切换到纯命令行模式(按Ctrl+Alt+F3进入tty),关闭桌面环境后再运行程序,避免Xorg等GUI进程占用GPU内存。
- 验证设备可用性:编译并运行NVIDIA CUDA示例中的
deviceQuery程序,检查GPU设备是否能被正常识别和初始化。 - 重启系统:若上述方法无效,重启系统清除驱动残留的内存泄漏问题,再重新测试程序。
内容的提问来源于stack exchange,提问作者Ustinian
相关产品推荐
相关产品推荐

