如何在TensorFlow C(++) API中实现Python版model() (__call__)针对小输入的推理性能?
Great question—this is a super common pain point when working with small batch sizes in TensorFlow, especially when moving between Python and C/C++. Let’s break down your questions and actionable solutions:
1. Achieving Python model() (call) Performance in TensorFlow C/C++ APIs
The core issue here is different inference execution paths:
- Python’s
model()uses TensorFlow 2.x’s eager execution, which is optimized for low-overhead, small-batch runs (it skips many graph-level optimizations that add unnecessary overhead for tiny inputs). predict()/predict_on_batch()and the traditionalTF_SessionRun()/model_bundle.GetSession()->Run()rely on the old graph execution pipeline, tuned for large batches with optimizations like graph fusion or auto-parallelism that backfire for small inputs.
To match Python’s eager performance in C/C++, you need to use TensorFlow’s eager execution C/C++ APIs (TFE for C, TF::Eager for C++):
C API Example (TFE)
#include <tensorflow/c/eager/c_api.h> // Initialize eager context with small-batch optimizations TF_Status* status = TF_NewStatus(); TFE_ContextOptions* ctx_opts = TFE_NewContextOptions(); TFE_ContextOptionsSetExperimentalOptimizeForSmallBatch(ctx_opts, true); TFE_Context* ctx = TFE_NewContext(ctx_opts, status); TFE_DeleteContextOptions(ctx_opts); // Load SavedModel into the eager context const char* saved_model_dir = "./<SavedModelDir>"; const char* tags[] = {"serve"}; TF_SavedModel* saved_model = TF_LoadSavedModel(saved_model_dir, tags, 1, status); TFE_SavedModel* eager_model = TFE_SavedModelImport(ctx, saved_model, status); // Fetch the serving function (match your model's signature name) TFE_SavedModelFunction* infer_func = TFE_SavedModelGetFunction(eager_model, "serving_default", status); // Prepare input tensor handles (replace with your actual input data) TFE_TensorHandle* inputs[] = { /* your input tensor handles */ }; int num_inputs = sizeof(inputs)/sizeof(TFE_TensorHandle*); // Execute inference (same low-overhead path as Python's model()) TFE_TensorHandle** outputs = TFE_Execute(infer_func, inputs, num_inputs, status); // Clean up resources TFE_DeleteTensorHandles(outputs, TFE_SavedModelFunctionGetNumOutputs(infer_func)); TFE_DeleteSavedModelFunction(infer_func); TFE_DeleteSavedModel(eager_model); TF_DeleteSavedModel(saved_model); TFE_DeleteContext(ctx); TF_DeleteStatus(status);
C++ API Example (TF::Eager)
For C++, use TensorFlow 2.x’s native eager API:
#include "tensorflow/core/eager/context.h" #include "tensorflow/core/saved_model/loader.h" // Initialize eager context tensorflow::Status status; std::unique_ptr<tensorflow::EagerContext> ctx = tensorflow::NewEagerContext( tensorflow::SessionOptions(), tensorflow::ContextDevicePlacementPolicy::DEVICE_PLACEMENT_EXPLICIT, &status); // Load SavedModel into eager mode tensorflow::SavedModelBundleLite bundle; status = tensorflow::LoadSavedModel( ctx.get(), tensorflow::SavedModelLoadOptions(), "./<SavedModelDir>", {"serve"}, &bundle); // Get the inference function const auto& func = bundle.GetFunction("serving_default"); // Prepare input tensors std::vector<tensorflow::Tensor> inputs = { /* your input tensors */ }; std::vector<tensorflow::Tensor> outputs; // Run inference via eager execution status = func->Run(inputs, &outputs);
Key notes:
- Enable small-batch optimizations in the eager context to mirror Python’s
model()behavior. - Skip large-batch focused optimizations (like auto-parallelism) to cut down on overhead.
2. Alternative APIs for Fast Small-Batch Inference
If you want even better performance than Python’s model(), TensorRT C++ API is the way to go—it’s purpose-built for GPU inference optimization, especially for small batches. Your earlier TensorRT SavedModel issue was likely due to precompiled TensorFlow binaries lacking TensorRT support; using TensorRT directly avoids this problem entirely:
Step 1: Convert SavedModel to TensorRT Engine
Use TensorRT’s Python API for easy conversion:
import tensorflow as tf from tensorflow.python.compiler.tensorrt import trt_convert as trt converter = trt.TrtGraphConverterV2(input_saved_model_dir="./<SavedModelDir>") # Optimize for your target small batch size converter.convert(max_workspace_size_bytes=1<<25, precision_mode="FP16") # Export to a TensorRT engine file (for C++ use) converter.save("./trt_engine_dir")
Step 2: Run Inference with TensorRT C++ API
#include <NvInfer.h> #include <fstream> // Minimal TensorRT logger class Logger : public nvinfer1::ILogger { void log(Severity severity, const char* msg) override { if (severity >= Severity::kWARNING) std::cout << msg << std::endl; } } gLogger; int main() { // Load pre-built TensorRT engine std::ifstream engine_file("model.engine", std::ios::binary); engine_file.seekg(0, std::ifstream::end); size_t engine_size = engine_file.tellg(); engine_file.seekg(0, std::ifstream::beg); char* engine_data = new char[engine_size]; engine_file.read(engine_data, engine_size); // Initialize TensorRT runtime and engine nvinfer1::IRuntime* runtime = nvinfer1::createInferRuntime(gLogger); nvinfer1::ICudaEngine* engine = runtime->deserializeCudaEngine(engine_data, engine_size, nullptr); delete[] engine_data; // Create execution context nvinfer1::IExecutionContext* context = engine->createExecutionContext(); // Prepare input/output buffers (adjust based on your model's layers) void* buffers[2]; int input_idx = engine->getBindingIndex("your_input_layer_name"); int output_idx = engine->getBindingIndex("your_output_layer_name"); cudaMalloc(&buffers[input_idx], 5 * input_size * sizeof(float)); // Batch size 5 cudaMalloc(&buffers[output_idx], 5 * output_size * sizeof(float)); // Copy input data from host to GPU cudaMemcpy(buffers[input_idx], host_input_data, 5 * input_size * sizeof(float), cudaMemcpyHostToDevice); // Run low-overhead inference context->executeV2(buffers); // Copy output back to host cudaMemcpy(host_output_data, buffers[output_idx], 5 * output_size * sizeof(float), cudaMemcpyDeviceToHost); // Cleanup cudaFree(buffers[input_idx]); cudaFree(buffers[output_idx]); context->destroy(); engine->destroy(); runtime->destroy(); return 0; }
TensorRT applies layer fusion, precision calibration (FP16/INT8), and batch-specific optimizations that far outperform TensorFlow’s default graph optimizations for small batches.
内容的提问来源于stack exchange,提问作者Olaf

