You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ONNXRuntime C++调用GPU时内存不足错误的修复问询

问题

在使用ONNXRuntime C++调用RealESRGAN_x4plus_anime_6B.onnx模型推理时,出现如下错误:

2024-04-21 10:24:38.313352771 [E:onnxruntime:, inference_session.cc:1798 operator()] Exception during initialization: /onnxruntime_src/onnxruntime/core/framework/bfc_arena.cc:376 void* onnxruntime::BFCArena::AllocateRawInternal(size_t, bool, onnxruntime::Stream*, bool, onnxruntime::WaitNotificationFn) Available memory of 0 is smaller than requested bytes of 256

原本想将GPU内存限制设为1GB,但代码里误写为cuda_options.gpu_mem_limit = 1 * 1024 * 1024(仅1MB),GPU有近2GB空闲内存,且ONNXRuntime Python版本能正常推理,询问修复方法。

使用的代码如下:

#include <iostream>
#include <onnxruntime_cxx_api.h>
#include "filesystem"
#include <opencv2/opencv.hpp>
#include <opencv2/dnn/dnn.hpp>

namespace fs = std::filesystem;


int main() {
    // Load the ONNX model
    fs::path d = fs::absolute(fs::path(__FILE__).parent_path());
    std::string onnx_model_path = (d / "RealESRGAN_x4plus_anime_6B.onnx").string();

    // Set up the ONNX Runtime session
    Ort::Env env(ORT_LOGGING_LEVEL_WARNING, "onnxruntime");
    Ort::SessionOptions session_options;
    session_options.SetGraphOptimizationLevel(GraphOptimizationLevel::ORT_ENABLE_EXTENDED);
    session_options.SetExecutionMode(ExecutionMode::ORT_SEQUENTIAL);
    OrtCUDAProviderOptions cuda_options;
    cuda_options.gpu_mem_limit = 1 * 1024 * 1024;

    session_options.AppendExecutionProvider_CUDA(cuda_options);
    Ort::Session ort_session(env, onnx_model_path.c_str(), session_options);

    // Load and preprocess input image
    std::string input_image_path = (d / "small.jpg").string();
    cv::Mat input_image = cv::imread(input_image_path);
    cv::cvtColor(input_image, input_image, cv::COLOR_BGR2RGB);
    input_image.convertTo(input_image, CV_32FC3, 1.0 / 255.0);
    cv::transpose(input_image, input_image);
    input_image = cv::dnn::blobFromImage(input_image);

    Ort::AllocatorWithDefaultOptions allocator;
    std::string input_name = ort_session.GetInputNameAllocated(0, allocator).get();
    std::string out_name = ort_session.GetOutputNameAllocated(0, allocator).get();

    // Perform inference
    std::vector<const char *> input_names = {input_name.c_str()};
    std::vector<const char *> out_names = {out_name.c_str()};

    Ort::MemoryInfo memoryInfo = Ort::MemoryInfo::CreateCpu(OrtAllocatorType::OrtArenaAllocator,
                                                            OrtMemType::OrtMemTypeDefault);

    // Get the shape of the input tensor
    std::vector<int64_t> input_shape = {1, input_image.channels(), input_image.rows, input_image.cols};

    // Allocate memory for the tensor data
    size_t tensor_size = input_image.total() * input_image.channels();
    float *tensor_data = new float[tensor_size];

// Copy the data from the cv::Mat to the tensor data buffer
    std::memcpy(tensor_data, input_image.data, tensor_size * sizeof(float));


    // Create the input tensor
    Ort::Value input_tensor = Ort::Value::CreateTensor<float>(memoryInfo, tensor_data,
                                                              input_image.total() * input_image.channels(),
                                                              input_shape.data(), input_shape.size());

    // Create a vector of input tensors
    // Create a vector of input tensors
    std::vector<Ort::Value> input_tensors = {};
    input_tensors.push_back(std::move(input_tensor)); // Move input_tensor into input_tensors
    std::vector<Ort::Value> out_tensors = ort_session.Run(Ort::RunOptions{nullptr}, input_names.data(),
                                                          input_tensors.data(), 1, out_names.data(), 1);

    // Post-process the output
    const std::vector<int64_t> out_shape = out_tensors[0].GetTensorTypeAndShapeInfo().GetShape();
    float *data_ptr = const_cast<float *>(out_tensors[0].GetTensorData<float>());
    cv::Mat out_img(out_shape[2], out_shape[3], CV_32FC3, data_ptr);

// Transpose the image
    cv::transpose(out_img, out_img);

// Scale the image to the range [0, 255]
    out_img = out_img * 255;

// Clamp the values to the range [0, 255]
    cv::max(0, out_img, out_img);
    cv::min(255, out_img, out_img);

// Convert the image to 8-bit unsigned integer type
    out_img.convertTo(out_img, CV_8UC3);

    // Save the output image
    std::string out_path = (d / "out.jpg").string();
    cv::Mat out_bgr;
    cv::cvtColor(out_img, out_bgr, cv::COLOR_RGB2BGR);
    cv::imwrite(out_path, out_bgr);
    std::cout << "Output saved to: " << out_path << std::endl;

    return 0;
}

修复方案

1. 修正GPU内存限制值

错误根源是cuda_options.gpu_mem_limit的单位是字节,你原本想设置1GB,但代码里只写了1MB(1*1024*1024),远不足以加载模型。将其修改为1GB对应的字节数:

cuda_options.gpu_mem_limit = 1LL * 1024 * 1024 * 1024;

这里用1LL确保计算时是64位整数,避免溢出。

2. 可选:不手动限制GPU内存(推荐)

如果GPU有足够空闲内存,完全可以不设置gpu_mem_limit,让ONNXRuntime自动按需分配内存,这样更灵活,也不会出现内存不足的问题:

OrtCUDAProviderOptions cuda_options;
// 移除 cuda_options.gpu_mem_limit = ... 这一行

修改后的核心代码片段

// Set up the ONNX Runtime session
Ort::Env env(ORT_LOGGING_LEVEL_WARNING, "onnxruntime");
Ort::SessionOptions session_options;
session_options.SetGraphOptimizationLevel(GraphOptimizationLevel::ORT_ENABLE_EXTENDED);
session_options.SetExecutionMode(ExecutionMode::ORT_SEQUENTIAL);
OrtCUDAProviderOptions cuda_options;
// 修正为1GB,或者直接删除下面这行
cuda_options.gpu_mem_limit = 1LL * 1024 * 1024 * 1024;

session_options.AppendExecutionProvider_CUDA(cuda_options);
Ort::Session ort_session(env, onnx_model_path.c_str(), session_options);

内容的提问来源于stack exchange,提问作者chikadance

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 09:57:33