TensorFlow Serving客户端性能优化困惑:修改张量生成为何耗时骤降
tf.contrib.util.make_tensor_proto with Manual TensorShapeProto Cuts Prediction Time Drastically? Great question—this massive speed difference makes total sense once you dig into what each approach is doing under the hood. Let’s break it down clearly:
tf.contrib.util.make_tensor_protois doing way more than you need
This utility function is built as a one-size-fits-all tool for converting Python/numpy data into TensorFlow protobufs. It automatically handles a ton of extra work that’s unnecessary for your fixed-shape use case:- Inferring tensor dimensions from raw input data (like scanning your image array to confirm its shape)
- Validating data types and running automatic conversions if needed
- Adding metadata to handle edge cases (sparse tensors, ragged arrays, etc.)
All these checks and auto-inference steps add up quickly, eating up 600+ ms of your prediction time when you already know exactly what your input shape should be.
Manual TensorShapeProto construction cuts out all the fluff
When you explicitly definedims = [tensor_shape_pb2.TensorShapeProto.Dim(size=1)]and build the proto by hand, you’re skipping every single auto-inference and validation step thatmake_tensor_protoruns. You’re telling the system exactly what the tensor shape is upfront—no guessing, no checks, no unnecessary conversions. This makes the protobuf serialization process lightning fast, and the resulting proto is far more compact than the one generated by the utility function.The serving side gets a simpler request to process
A manually constructed proto also reduces work on the TensorFlow Serving end. The server doesn’t have to parse and validate an inferred shape from the proto; it can immediately map the incoming data to the model’s input tensor. This cuts down on deserialization and setup time on the server, contributing to that huge speed jump from 600ms to 20ms.
One important caveat: this optimization only works because you’re 100% sure of your input shape and data type. If your input shapes vary or you need to handle dynamic data, make_tensor_proto’s safety checks are worth the overhead to avoid errors. But for fixed-shape, predictable requests like your image inference, manual proto construction is a massive win.
内容的提问来源于stack exchange,提问作者tianxiang fei

