You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ONNX进行TRT推理报错:Error Code 1: Cuda Driver (invalid resource handle)

解决TensorRT推理时"Error Code 1: Cuda Driver (invalid resource handle)"错误

以下是你的代码中导致错误的核心问题及修复方案:

核心错误点

  1. 内存大小计算错误
    代码中size = trt.volume(engine.get_binding_shape(binding)) * -1的*-1是完全错误的操作,trt.volume()返回的是当前绑定维度的总元素数量,负数会导致分配的内存空间异常,直接触发CUDA资源句柄无效的错误。

  2. 输入数据类型不匹配
    你预处理后的输入是float32类型,但代码中强制指定trt_types = [trt.int32],类型不匹配会导致内存拷贝和推理阶段的资源访问错误。

  3. CUDA上下文管理优化
    建议在创建TensorRT Runtime和Engine之前就激活CUDA上下文,避免上下文切换导致的资源问题。

修复后的validate_trt_result函数

def validate_trt_result(self, input_path):
    TRT_LOGGER = trt.Logger(trt.Logger.VERBOSE)
    
    trt_file_name = "PATH_TO_TRT_FILE"

    # 先初始化CUDA上下文
    cuda.init()
    device = cuda.Device(0)
    ctx = device.make_context()

    trt_runtime = trt.Runtime(TRT_LOGGER)

    with open(trt_file_name, 'rb') as f:
        engine_data = f.read()

    engine = trt_runtime.deserialize_cuda_engine(engine_data)

    inputs, outputs, bindings = [], [], []

    context = engine.create_execution_context()
    stream = cuda.Stream()
    
    index = 0
    for binding in engine:
        # 修复:去掉*-1,正确计算元素数量
        binding_shape = engine.get_binding_shape(binding)
        # 如果是动态shape,使用设置后的shape计算
        if context.get_binding_shape(index) != tuple(binding_shape):
            binding_shape = context.get_binding_shape(index)
        size = trt.volume(binding_shape)
        dtype = trt.nptype(engine.get_binding_dtype(binding))
        host_mem = cuda.pagelocked_empty(size, dtype)
        device_mem = cuda.mem_alloc(host_mem.nbytes)
        bindings.append(int(device_mem))
        if engine.binding_is_input(binding):
            inputs.append(HostDeviceMem(host_mem, device_mem))
            # 设置实际推理的batch shape
            context.set_binding_shape(index, [1, 3, IMG_SIZE, IMG_SIZE])
        else:
            outputs.append(HostDeviceMem(host_mem, device_mem))
        index += 1

    # 确认所有绑定shape都已指定
    assert context.all_binding_shapes_specified, "Not all binding shapes are specified!"

    # 输入预处理
    input_img = cv2.imread(input_path)
    input_r = cv2.resize(input_img, dsize=(256, 256))
    input_p = np.transpose(input_r, (2, 0, 1))  
    input_e = np.expand_dims(input_p, axis=0)
    input_f = input_e.astype(np.float32)
    input_f /= 255      
    
    numpy_array_input = [input_f]
    hosts = [input.host for input in inputs]
    # 修复:使用模型输入的实际dtype,而不是硬编码int32
    trt_types = [engine.get_binding_dtype(binding) for binding in engine if engine.binding_is_input(binding)]
    
    for numpy_array, host, trt_type in zip(numpy_array_input, hosts, trt_types):
        numpy_array = np.asarray(numpy_array).astype(trt.nptype(trt_type)).ravel()
        print(numpy_array.shape)
        np.copyto(host, numpy_array)

    # 异步内存拷贝
    [cuda.memcpy_htod_async(inp.device, inp.host, stream) for inp in inputs]

    # 执行推理
    context.execute_async_v2(bindings=bindings, stream_handle=stream.handle)

    # 拷贝结果回主机
    [cuda.memcpy_dtoh_async(out.host, out.device, stream) for out in outputs]
    stream.synchronize()

    print("TRT model inference result : ")
    output = outputs[0].host
    for one in output :
        print(one)
    
    # 清理上下文和资源
    ctx.pop()
    del context, engine, trt_runtime, stream

额外注意事项

  • 确保HostDeviceMem类的定义正确,通常应该是包含host和device属性的简单类:
    class HostDeviceMem:
        def __init__(self, host_mem, device_mem):
            self.host = host_mem
            self.device = device_mem
    
  • 如果你在trt_export中启用了FP16模式,确保推理时输入数据的 dtype 和模型一致(如果模型是FP16,输入需要转为float16)。
  • 运行代码前确认CUDA驱动版本和TensorRT版本兼容,避免版本不匹配导致的底层资源错误。

内容的提问来源于stack exchange,提问作者happychild

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 19:04:07