You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用cudaSetDevice()出现"out of memory"错误的原因是什么?

关于cudaSetDevice()触发"out of memory"错误的问题

我编写了一个C++程序,用于查询GPU设备数量、获取第一台设备属性后调用cudaSetDevice()。程序在虚拟机编译完成后,拷贝到配备NVIDIA RTX PRO 5000 Blackwell GPU的同配置物理机运行,输出如下:

Found 1 GPUs, trying to work with the first one (id 0).
GPU name: NVIDIA RTX PRO 5000 Blackwell
Trying to cudaSetDevice(0): cudaSetDevice failed, with error: out of memory

已知前序API调用(cudaGetDeviceCount、cudaGetDeviceProperties)均成功,无遗留错误,现咨询:为何会出现此错误?cudaSetDevice()为何会触发“out of memory”错误?

机器配置信息

  • 操作系统发行版:SLES 15 SP6
  • 内核:Linux 6.4.0-150600.23.17-default
  • CUDA:版本12.9.1,路径/usr/local/cuda,NVCC版本V12.9.86
  • CPU架构:Intel x86_64

差异点

  • 编译环境:虚拟机(无物理GPU)
  • 运行环境:配备RTX PRO 5000 Blackwell,驱动版本580.105.08(高于CUDA 12.9自带的驱动版本)

测试代码

#include <cuda_runtime.h>
#include <iostream>

int main(int argc, char **argv)
{
    int device_count = 0;
    cudaError_t status;
    status = cudaGetDeviceCount(&device_count);
    if (status != cudaSuccess) {
        std::cout << "cudaGetDeviceCount failed: " << cudaGetErrorString(status) << '\n';
        return -1;
    }
    if (device_count == 0) {
        std::cout << "No GPUs found." << std::endl;
        return 0;
    }
    std::cout << "Found " << device_count << " GPUs, trying to work with the first one (id 0).\n";
    int device_id = 0;
    cudaDeviceProp device_prop;
    status = cudaGetDeviceProperties(&device_prop, device_id);
    if (status != cudaSuccess) {
        std::cout << "Failed to obtain device properties\n";
        return -1;
    }
    std::cout << "GPU name: " << device_prop.name << '\n';
    std::cout << "Trying to cudaSetDevice(" << device_id << "): ";
    status = cudaSetDevice(device_id);
    if (status != cudaSuccess) {
        std::cout << "cudaSetDevice failed, with error: " << cudaGetErrorString(status) << '\n';
        return -1;
    }
}

问题分析与可能原因

cudaSetDevice()返回out of memory错误,通常并非指GPU显存不足,而是CUDA Runtime在初始化设备上下文时无法分配必要资源,常见原因包括:

  1. 驱动与CUDA版本兼容性bug
    虽然驱动版本高于CUDA 12.9自带版本,但Blackwell架构对CUDA和驱动的组合有特定要求。若两者间存在兼容性问题,可能导致上下文初始化失败,错误被误报为内存不足。

  2. 跨环境编译的二进制兼容性问题
    在无GPU的虚拟机编译时,NVCC可能默认未生成针对Blackwell架构(sm_90)的PTX或二进制代码,导致运行时无法加载适配目标GPU的代码,进而触发资源分配类错误。

  3. GPU关键资源被抢占
    运行机器的GPU可能被其他进程(如桌面渲染服务、其他CUDA程序)占用了上下文初始化所需的核心资源,导致cudaSetDevice()无法完成创建流程。

  4. 系统安全模块限制
    SLES 15 SP6默认启用的SELinux/AppArmor可能限制了程序对GPU设备文件的访问权限,导致初始化时无法分配系统资源,错误表现为内存不足。

排查建议

  • 验证版本兼容性:查阅NVIDIA官方文档确认驱动580.105.08与CUDA 12.9.1的适配性;尝试降级到CUDA 12.9推荐的驱动版本(如555.x系列)测试。
  • 重新编译程序:在目标GPU机器上重新编译,添加明确的架构编译选项:nvcc -arch=sm_90 your_code.cpp -o your_program,确保生成适配Blackwell的二进制代码。
  • 检查GPU资源占用:执行nvidia-smi查看GPU进程列表和内存使用情况,关闭非必要占用进程后重新运行程序。
  • 排查安全限制:临时关闭SELinux(setenforce 0)或调整AppArmor规则,验证是否为权限问题导致的错误。

内容的提问来源于stack exchange,提问作者einpoklum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.11 14:04:51