You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ONNX Runtime I/O Binding绑定后输出张量返回空指针求助

ONNX Runtime I/O Binding输出张量返回NULL指针排查建议

使用ONNX Runtime的I/O Binding绑定张量输入与输出时,输出张量返回NULL指针,已核对输入、输出张量的数据与形状,问题仍存在,代码如下:

std::vector<Ort::Value> input_tensors;
std::vector<Ort::Value> output_tensors;
std::vector<const char*> input_node_names_c_str;
std::vector<const char*> output_node_names_c_str;
int64_t input_height = input_node_dims[0].at(2);
int64_t input_width = input_node_dims[0].at(3);

// Pass gpu_graph_id to RunOptions through RunConfigs
Ort::RunOptions run_option;
// gpu_graph_id is optional if the session uses only one cuda graph
run_option.AddConfigEntry("gpu_graph_id", "1");

// Dimension expansion [CHW -> NCHW]
std::vector<int64_t> input_tensor_shape = {1, 3, input_height, input_width};
std::vector<int64_t> output_tensor_shape = {1, 300, 84};
size_t input_tensor_size = vector_product(input_tensor_shape);
size_t output_tensor_size = vector_product(output_tensor_shape);
std::vector<float> input_tensor_values(p_blob, p_blob + input_tensor_size);

Ort::IoBinding io_binding{session};
Ort::MemoryInfo memory_info = Ort::MemoryInfo::CreateCpu(OrtDeviceAllocator, OrtMemTypeCPU);

input_tensors.push_back(Ort::Value::CreateTensor<float>(
        memory_info, input_tensor_values.data(), input_tensor_size,
        input_tensor_shape.data(), input_tensor_shape.size()
));

// Check if input and output node names are empty
for (const auto& inputNodeName : input_node_names) {
    if (std::string(inputNodeName).empty()) {
        std::cerr << "Empty input node name found." << std::endl;
    }
}

// format conversion
for (const auto& inputName : input_node_names) {
    input_node_names_c_str.push_back(inputName.c_str());
}

for (const auto& outputName : output_node_names) {
    output_node_names_c_str.push_back(outputName.c_str());
}

io_binding.BindInput(input_node_names_c_str[0], input_tensors[0]);

Ort::MemoryInfo output_mem_info{"Cuda", OrtDeviceAllocator, 0,
                                OrtMemTypeDefault};

cudaMalloc(&output_data_ptr, output_tensor_size * sizeof(float));
output_tensors.push_back(Ort::Value::CreateTensor<float>(
    output_mem_info,  static_cast<float*>(output_data_ptr),output_tensor_size, 
    output_tensor_shape.data(),output_tensor_shape.size()));                            

io_binding.BindOutput(output_node_names_c_str[0],  output_tensors[0]);
session.Run(run_option, io_binding);

//Get output results
auto* rawOutput = output_tensors[0].GetTensorData<float>();
cout<<rawOutput<<endl; //suhail
cudaFree(output_data_ptr); //suhail
std::vector<int64_t> outputShape = output_tensors[0].GetTensorTypeAndShapeInfo().GetShape();
for(auto i:outputShape){cout<<i<<" ";} cout<<endl; //suhail
size_t count = output_tensors[0].GetTensorTypeAndShapeInfo().GetElementCount();
cout<<count<<endl; //suhail
std::vector<float> output(rawOutput, rawOutput + count);

排查解决建议

  • 检查CUDA内存分配有效性:在cudaMalloc之后立即验证分配是否成功,CUDA内存分配失败会直接导致后续张量的数据指针为空。添加判断:

    if (cudaMalloc(&output_data_ptr, output_tensor_size * sizeof(float)) != cudaSuccess || output_data_ptr == nullptr) {
        std::cerr << "CUDA memory allocation failed" << std::endl;
        return;
    }
    
  • 确认Session设备与输出内存匹配:如果Session基于CPU创建,绑定CUDA内存作为输出会不兼容。确保Session初始化时指定了CUDA执行提供者,比如通过Ort::SessionOptions配置CUDA EP。

  • 添加CUDA同步操作:使用CUDA图或异步执行时,session.Run()可能不会立即完成计算,导致访问输出张量时数据未写入。在session.Run()之后添加:

    cudaDeviceSynchronize();
    

    确保设备计算完成后再读取数据。

  • 修正内存释放时机:代码中在获取rawOutput后立即调用cudaFree(output_data_ptr),随后用已释放的指针构造std::vector<float>属于未定义行为。调整顺序:

    //Get output results
    auto* rawOutput = output_tensors[0].GetTensorData<float>();
    cout<<rawOutput<<endl; 
    std::vector<int64_t> outputShape = output_tensors[0].GetTensorTypeAndShapeInfo().GetShape();
    for(auto i:outputShape){cout<<i<<" ";} cout<<endl; 
    size_t count = output_tensors[0].GetTensorTypeAndShapeInfo().GetElementCount();
    cout<<count<<endl; 
    std::vector<float> output(rawOutput, rawOutput + count);
    cudaFree(output_data_ptr); // 移到构造vector之后
    
  • 验证输出节点名称准确性:确保output_node_names_c_str[0]与ONNX模型实际输出节点名称完全一致(区分大小写),可使用Netron工具查看模型的输入输出节点名称。

  • 启用ONNX Runtime日志排查:设置RunOptions的日志级别,查看运行时是否有错误或警告信息:

    run_option.SetLogLevel(ORT_LOG_LEVEL_VERBOSE);
    run_option.SetLogFunction([](OrtLogLevel level, const char* log_message) {
        std::cerr << "ONNX Runtime Log: " << log_message << std::endl;
    });
    

内容的提问来源于stack exchange,提问作者Suhail Muhammed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 19:02:08