TensorFlow C++ API批量推理问题:批次大于1时仅首样本结果正确
Hey there, let's troubleshoot why only your first sample returns correct results during batch inference with TensorFlow's C++ API. I’ve run into similar quirks before, so here are the key areas to check and fix:
First off, your input node main_input has a shape of [-1, 8, 8, 13] (the -1 lets TensorFlow handle dynamic batch sizes). In C++, you need to explicitly define the batch dimension when creating your input tensor, and make sure you populate data for all samples correctly:
- When initializing the tensor, set the shape to match your batch size:
int batch_size = 3; // Replace with your actual batch size Tensor input_tensor(DT_FLOAT, TensorShape({batch_size, 8, 8, 13})); - Avoid common data population mistakes: Don’t just fill the first
8*8*13elements and leave the rest uninitialized. Use the tensor’s flat data pointer to iterate through all samples, maintaining the[batch, height, width, channels]order:auto input_ptr = input_tensor.flat<float>().data(); for (int b = 0; b < batch_size; ++b) { for (int h = 0; h < 8; ++h) { for (int w = 0; w < 8; ++w) { for (int c = 0; c < 13; ++c) { // Replace with your actual sample data access input_ptr[b*8*8*13 + h*8*13 + w*13 + c] = your_batch_data[b][h][w][c]; } } } }
It’s easy to pick the wrong output node after converting from Keras to .pb:
- Use a tool like TensorBoard or
saved_model_clito inspect your.pbmodel’s nodes. Your output node should have a shape of[-1, 1](not[1]) to support batch outputs. For example, if your Keras model ends with aDense(1)layer, the corresponding TensorFlow node might be named something likedense_1/BiasAddorIdentity. - If you pick a node that only outputs a single value (shape
[1]), it’ll only return the first sample’s result every time.
Before debugging C++ code, rule out model conversion issues:
- Load the
.pbmodel in Python and run a batch inference test with the same input data you’re using in C++. If the Python test returns correct results for all samples, the problem is in your C++ implementation. If Python also fails, re-run the keras2tensorflow conversion with explicit input/output node flags, like:# Example conversion command (adjust node names to match your model) keras2tensorflow --input_model your_keras_model.h5 --output_model output.pb --input_nodes main_input --output_nodes your_output_node
In your C++ inference code, make sure you’re fetching the output correctly and not accidentally slicing only the first element:
- When calling
session->Run(), your output tensor will have a shape of[batch_size, 1]. Access each sample’s result by indexing into the flat tensor:std::vector<std::pair<string, Tensor>> inputs = {{"main_input", input_tensor}}; std::vector<Tensor> outputs; Status status = session->Run(inputs, {"your_output_node_name"}, {}, &outputs); if (status.ok()) { auto output_ptr = outputs[0].flat<float>().data(); for (int b = 0; b < batch_size; ++b) { float result = output_ptr[b]; // Process each sample's result here } }
内容的提问来源于stack exchange,提问作者danny

