You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用TensorRT进行INT8推理时输出数据量异常增大的技术咨询

问题分析与解决方案

首先纠正一个小细节:你提到INT8输出大小是预期的5倍,但实际计算下来 13680000 / 273600 = 50,刚好是你训练时的batch size。结合FP32/FP16推理输出正常的情况,核心问题应该是你的INT8引擎被固定为batch size 50的输出维度,而FP32/FP16引擎是按batch size 1构建的。

下面分几种可能的原因和对应的解决办法:

1. 引擎构建时未启用动态Batch Size,输入Shape固定为训练时的50

如果构建INT8引擎时,你没有设置动态Shape,而是直接用训练时的batch size 50作为输入维度,引擎会被固化为只能处理batch size 50的输入,输出自然也是50个样本的结果。而你FP32/FP16的引擎可能是按batch size 1构建的,所以输出正常。

解决办法:

构建INT8引擎时启用动态Shape,通过TensorRT的优化配置文件指定支持的batch size范围。示例代码如下:

import tensorrt as trt

TRT_LOGGER = trt.Logger(trt.Logger.WARNING)
builder = trt.Builder(TRT_LOGGER)
# 必须启用EXPLICIT_BATCH才能支持动态Shape
network = builder.create_network(1 << int(trt.NetworkDefinitionCreationFlag.EXPLICIT_BATCH))
config = builder.create_builder_config()

# 创建优化配置文件,设置输入的动态Shape(这里假设输入是(3, 60, 80),batch维度设为1)
profile = builder.create_optimization_profile()
input_name = network.get_input(0).name
# 参数格式:(输入名, 最小Shape, 最优Shape, 最大Shape)
profile.set_shape(input_name, (1, 3, 60, 80), (1, 3, 60, 80), (1, 3, 60, 80))
config.add_optimization_profile(profile)

# 配置INT8校准
config.set_flag(trt.BuilderFlag.INT8)
config.int8_calibrator = YourCalibrator()  # 替换成你自己的校准器类

# 构建引擎
engine = builder.build_engine(network, config)

2. 推理时未设置Context的Binding Shape

即使引擎支持动态Shape,如果你在推理时没有明确设置当前使用的batch size,Context可能会默认使用校准或构建时的batch size(也就是50),导致输出维度还是50。

解决办法:

在创建Context后、分配缓冲区前,手动设置输入的Binding Shape为batch size 1:

with engine.create_execution_context() as context:
    # 设置输入的Shape为batch 1(替换成你的实际输入Shape)
    context.set_binding_shape(0, (1, 3, 60, 80))
    # 打印确认输入输出的Shape是否正确
    print(f"Input binding shape: {context.get_binding_shape(0)}")
    print(f"Output binding shape: {context.get_binding_shape(1)}")
    
    fps_time = time.time()
    inputs, outputs, bindings, stream = common.allocate_buffers(engine, context)
    im = np.array(frm, dtype=np.float32, order='C')
    inputs[0].host = im.flatten()
    [outputs] = common.do_inference(context, bindings=bindings, inputs=inputs, outputs=outputs, stream=stream, batch_size=1)
    outputs = outputs.reshape((60, 80, 57))

3. 旧版common.allocate_buffers函数的兼容性问题

如果你使用的是TensorRT示例中的common.py,某些旧版本的allocate_buffers函数可能没有正确处理动态Shape或INT8张量的维度,导致缓冲区分配的大小还是固定的batch 50。

解决办法:

修改allocate_buffers函数,确保它根据Context的实际Binding Shape来计算缓冲区大小,而不是依赖引擎的静态Shape。比如,在函数中加入获取Context Shape的逻辑:

def allocate_buffers(engine, context=None):
    inputs = []
    outputs = []
    bindings = []
    stream = cuda.Stream()
    for binding in engine:
        # 如果有Context,用Context的Binding Shape,否则用引擎的静态Shape
        shape = context.get_binding_shape(binding) if context else engine.get_binding_shape(binding)
        size = trt.volume(shape)
        dtype = trt.nptype(engine.get_binding_dtype(binding))
        # 分配主机和设备缓冲区
        host_mem = cuda.pagelocked_empty(size, dtype)
        device_mem = cuda.mem_alloc(host_mem.nbytes)
        # 将缓冲区绑定到引擎的binding
        bindings.append(int(device_mem))
        if engine.binding_is_input(binding):
            inputs.append(HostDeviceMem(host_mem, device_mem))
        else:
            outputs.append(HostDeviceMem(host_mem, device_mem))
    return inputs, outputs, bindings, stream

快速验证方法

在你的推理代码中加入一行打印输出Shape的代码,就能确认是不是batch维度的问题:

with engine.create_execution_context() as context:
    print(f"Expected output shape: (1, 60, 80, 57)")
    print(f"Actual output shape: {context.get_binding_shape(1)}")

如果实际输出Shape是(50, 60, 80, 57),那就能坐实是batch size被固定的问题了。

内容的提问来源于stack exchange,提问作者batuman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:29:55