使用ONNXRuntime C++调用GPU时内存不足错误的修复问询
问题
在使用ONNXRuntime C++调用RealESRGAN_x4plus_anime_6B.onnx模型推理时,出现如下错误:
2024-04-21 10:24:38.313352771 [E:onnxruntime:, inference_session.cc:1798 operator()] Exception during initialization: /onnxruntime_src/onnxruntime/core/framework/bfc_arena.cc:376 void* onnxruntime::BFCArena::AllocateRawInternal(size_t, bool, onnxruntime::Stream*, bool, onnxruntime::WaitNotificationFn) Available memory of 0 is smaller than requested bytes of 256
原本想将GPU内存限制设为1GB,但代码里误写为cuda_options.gpu_mem_limit = 1 * 1024 * 1024(仅1MB),GPU有近2GB空闲内存,且ONNXRuntime Python版本能正常推理,询问修复方法。
使用的代码如下:
#include <iostream> #include <onnxruntime_cxx_api.h> #include "filesystem" #include <opencv2/opencv.hpp> #include <opencv2/dnn/dnn.hpp> namespace fs = std::filesystem; int main() { // Load the ONNX model fs::path d = fs::absolute(fs::path(__FILE__).parent_path()); std::string onnx_model_path = (d / "RealESRGAN_x4plus_anime_6B.onnx").string(); // Set up the ONNX Runtime session Ort::Env env(ORT_LOGGING_LEVEL_WARNING, "onnxruntime"); Ort::SessionOptions session_options; session_options.SetGraphOptimizationLevel(GraphOptimizationLevel::ORT_ENABLE_EXTENDED); session_options.SetExecutionMode(ExecutionMode::ORT_SEQUENTIAL); OrtCUDAProviderOptions cuda_options; cuda_options.gpu_mem_limit = 1 * 1024 * 1024; session_options.AppendExecutionProvider_CUDA(cuda_options); Ort::Session ort_session(env, onnx_model_path.c_str(), session_options); // Load and preprocess input image std::string input_image_path = (d / "small.jpg").string(); cv::Mat input_image = cv::imread(input_image_path); cv::cvtColor(input_image, input_image, cv::COLOR_BGR2RGB); input_image.convertTo(input_image, CV_32FC3, 1.0 / 255.0); cv::transpose(input_image, input_image); input_image = cv::dnn::blobFromImage(input_image); Ort::AllocatorWithDefaultOptions allocator; std::string input_name = ort_session.GetInputNameAllocated(0, allocator).get(); std::string out_name = ort_session.GetOutputNameAllocated(0, allocator).get(); // Perform inference std::vector<const char *> input_names = {input_name.c_str()}; std::vector<const char *> out_names = {out_name.c_str()}; Ort::MemoryInfo memoryInfo = Ort::MemoryInfo::CreateCpu(OrtAllocatorType::OrtArenaAllocator, OrtMemType::OrtMemTypeDefault); // Get the shape of the input tensor std::vector<int64_t> input_shape = {1, input_image.channels(), input_image.rows, input_image.cols}; // Allocate memory for the tensor data size_t tensor_size = input_image.total() * input_image.channels(); float *tensor_data = new float[tensor_size]; // Copy the data from the cv::Mat to the tensor data buffer std::memcpy(tensor_data, input_image.data, tensor_size * sizeof(float)); // Create the input tensor Ort::Value input_tensor = Ort::Value::CreateTensor<float>(memoryInfo, tensor_data, input_image.total() * input_image.channels(), input_shape.data(), input_shape.size()); // Create a vector of input tensors // Create a vector of input tensors std::vector<Ort::Value> input_tensors = {}; input_tensors.push_back(std::move(input_tensor)); // Move input_tensor into input_tensors std::vector<Ort::Value> out_tensors = ort_session.Run(Ort::RunOptions{nullptr}, input_names.data(), input_tensors.data(), 1, out_names.data(), 1); // Post-process the output const std::vector<int64_t> out_shape = out_tensors[0].GetTensorTypeAndShapeInfo().GetShape(); float *data_ptr = const_cast<float *>(out_tensors[0].GetTensorData<float>()); cv::Mat out_img(out_shape[2], out_shape[3], CV_32FC3, data_ptr); // Transpose the image cv::transpose(out_img, out_img); // Scale the image to the range [0, 255] out_img = out_img * 255; // Clamp the values to the range [0, 255] cv::max(0, out_img, out_img); cv::min(255, out_img, out_img); // Convert the image to 8-bit unsigned integer type out_img.convertTo(out_img, CV_8UC3); // Save the output image std::string out_path = (d / "out.jpg").string(); cv::Mat out_bgr; cv::cvtColor(out_img, out_bgr, cv::COLOR_RGB2BGR); cv::imwrite(out_path, out_bgr); std::cout << "Output saved to: " << out_path << std::endl; return 0; }
修复方案
1. 修正GPU内存限制值
错误根源是cuda_options.gpu_mem_limit的单位是字节,你原本想设置1GB,但代码里只写了1MB(1*1024*1024),远不足以加载模型。将其修改为1GB对应的字节数:
cuda_options.gpu_mem_limit = 1LL * 1024 * 1024 * 1024;
这里用1LL确保计算时是64位整数,避免溢出。
2. 可选:不手动限制GPU内存(推荐)
如果GPU有足够空闲内存,完全可以不设置gpu_mem_limit,让ONNXRuntime自动按需分配内存,这样更灵活,也不会出现内存不足的问题:
OrtCUDAProviderOptions cuda_options; // 移除 cuda_options.gpu_mem_limit = ... 这一行
修改后的核心代码片段
// Set up the ONNX Runtime session Ort::Env env(ORT_LOGGING_LEVEL_WARNING, "onnxruntime"); Ort::SessionOptions session_options; session_options.SetGraphOptimizationLevel(GraphOptimizationLevel::ORT_ENABLE_EXTENDED); session_options.SetExecutionMode(ExecutionMode::ORT_SEQUENTIAL); OrtCUDAProviderOptions cuda_options; // 修正为1GB,或者直接删除下面这行 cuda_options.gpu_mem_limit = 1LL * 1024 * 1024 * 1024; session_options.AppendExecutionProvider_CUDA(cuda_options); Ort::Session ort_session(env, onnx_model_path.c_str(), session_options);
内容的提问来源于stack exchange,提问作者chikadance
相关产品推荐
相关产品推荐

