使用TensorRT进行INT8推理时输出数据量异常增大的技术咨询
首先纠正一个小细节:你提到INT8输出大小是预期的5倍,但实际计算下来 13680000 / 273600 = 50,刚好是你训练时的batch size。结合FP32/FP16推理输出正常的情况,核心问题应该是你的INT8引擎被固定为batch size 50的输出维度,而FP32/FP16引擎是按batch size 1构建的。
下面分几种可能的原因和对应的解决办法:
1. 引擎构建时未启用动态Batch Size,输入Shape固定为训练时的50
如果构建INT8引擎时,你没有设置动态Shape,而是直接用训练时的batch size 50作为输入维度,引擎会被固化为只能处理batch size 50的输入,输出自然也是50个样本的结果。而你FP32/FP16的引擎可能是按batch size 1构建的,所以输出正常。
解决办法:
构建INT8引擎时启用动态Shape,通过TensorRT的优化配置文件指定支持的batch size范围。示例代码如下:
import tensorrt as trt TRT_LOGGER = trt.Logger(trt.Logger.WARNING) builder = trt.Builder(TRT_LOGGER) # 必须启用EXPLICIT_BATCH才能支持动态Shape network = builder.create_network(1 << int(trt.NetworkDefinitionCreationFlag.EXPLICIT_BATCH)) config = builder.create_builder_config() # 创建优化配置文件,设置输入的动态Shape(这里假设输入是(3, 60, 80),batch维度设为1) profile = builder.create_optimization_profile() input_name = network.get_input(0).name # 参数格式:(输入名, 最小Shape, 最优Shape, 最大Shape) profile.set_shape(input_name, (1, 3, 60, 80), (1, 3, 60, 80), (1, 3, 60, 80)) config.add_optimization_profile(profile) # 配置INT8校准 config.set_flag(trt.BuilderFlag.INT8) config.int8_calibrator = YourCalibrator() # 替换成你自己的校准器类 # 构建引擎 engine = builder.build_engine(network, config)
2. 推理时未设置Context的Binding Shape
即使引擎支持动态Shape,如果你在推理时没有明确设置当前使用的batch size,Context可能会默认使用校准或构建时的batch size(也就是50),导致输出维度还是50。
解决办法:
在创建Context后、分配缓冲区前,手动设置输入的Binding Shape为batch size 1:
with engine.create_execution_context() as context: # 设置输入的Shape为batch 1(替换成你的实际输入Shape) context.set_binding_shape(0, (1, 3, 60, 80)) # 打印确认输入输出的Shape是否正确 print(f"Input binding shape: {context.get_binding_shape(0)}") print(f"Output binding shape: {context.get_binding_shape(1)}") fps_time = time.time() inputs, outputs, bindings, stream = common.allocate_buffers(engine, context) im = np.array(frm, dtype=np.float32, order='C') inputs[0].host = im.flatten() [outputs] = common.do_inference(context, bindings=bindings, inputs=inputs, outputs=outputs, stream=stream, batch_size=1) outputs = outputs.reshape((60, 80, 57))
3. 旧版common.allocate_buffers函数的兼容性问题
如果你使用的是TensorRT示例中的common.py,某些旧版本的allocate_buffers函数可能没有正确处理动态Shape或INT8张量的维度,导致缓冲区分配的大小还是固定的batch 50。
解决办法:
修改allocate_buffers函数,确保它根据Context的实际Binding Shape来计算缓冲区大小,而不是依赖引擎的静态Shape。比如,在函数中加入获取Context Shape的逻辑:
def allocate_buffers(engine, context=None): inputs = [] outputs = [] bindings = [] stream = cuda.Stream() for binding in engine: # 如果有Context,用Context的Binding Shape,否则用引擎的静态Shape shape = context.get_binding_shape(binding) if context else engine.get_binding_shape(binding) size = trt.volume(shape) dtype = trt.nptype(engine.get_binding_dtype(binding)) # 分配主机和设备缓冲区 host_mem = cuda.pagelocked_empty(size, dtype) device_mem = cuda.mem_alloc(host_mem.nbytes) # 将缓冲区绑定到引擎的binding bindings.append(int(device_mem)) if engine.binding_is_input(binding): inputs.append(HostDeviceMem(host_mem, device_mem)) else: outputs.append(HostDeviceMem(host_mem, device_mem)) return inputs, outputs, bindings, stream
快速验证方法
在你的推理代码中加入一行打印输出Shape的代码,就能确认是不是batch维度的问题:
with engine.create_execution_context() as context: print(f"Expected output shape: (1, 60, 80, 57)") print(f"Actual output shape: {context.get_binding_shape(1)}")
如果实际输出Shape是(50, 60, 80, 57),那就能坐实是batch size被固定的问题了。
内容的提问来源于stack exchange,提问作者batuman

