You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用TensorRT加速PyTorch模型推理速度未提升,求原因排查

TensorRT加速PyTorch模型无效的问题排查

我尝试用TensorRT提升PyTorch模型的推理速度,写了一段测试代码,但使用TensorRT和不使用的FPS都是约14.2,没有速度提升。请问问题出在哪?有没有遗漏的配置?

测试代码

import torch
import argparse
import time
import numpy as np
import torch_tensorrt

# Define a simple PyTorch model
class MyModel(torch.nn.Module):
    def __init__(self):
        super().__init__()
        self.conv1 = torch.nn.Conv2d(3, 32, kernel_size=3, stride=1, padding=1)
        self.relu1 = torch.nn.ReLU()
        self.conv2 = torch.nn.Conv2d(32, 64, kernel_size=3, stride=1, padding=1)
        self.relu2 = torch.nn.ReLU()
        self.pool = torch.nn.MaxPool2d(kernel_size=2, stride=2)
        self.fc1 = torch.nn.Linear(64 * 16 * 16, 512)
        self.relu3 = torch.nn.ReLU()
        self.fc2 = torch.nn.Linear(512, 10)

    def forward(self, x):
        x = self.conv1(x)
        x = self.relu1(x)
        x = self.conv2(x)
        x = self.relu2(x)
        x = self.pool(x)
        x = x.view(-1, 64 * 16 * 16)
        x = self.fc1(x)
        x = self.relu3(x)
        x = self.fc2(x)
        return x

def compute(use_tensorrt=False):
    force_cpu = False
    useCuda = torch.cuda.is_available() and not force_cpu
    if useCuda:
        print('Using CUDA.')
        dtype = torch.cuda.FloatTensor
        ltype = torch.cuda.LongTensor
        device = torch.device("cuda:0")
    else:
        print('No CUDA available.')
        dtype = torch.FloatTensor
        ltype = torch.LongTensor
        device = torch.device("cpu")

    model = MyModel()

    input_shape = (8192, 3, 32, 32)

    if use_tensorrt:
        model = torch.compile(
            model,
            backend="torch_tensorrt",
            options={
                "truncate_long_and_double": True,
                "precision": dtype,
                "workspace_size" : 20 << 30
            },
            dynamic=False,
        )

    model = model.to(device)
    model.eval()

    num_iterations = 100
    total_time = 0.0
    with torch.no_grad():
        input_data = torch.randn(input_shape).to(device).type(dtype)
        #warmup
        for i in range(100):
            output_data = model(input_data)

        for i in range(num_iterations):
            start_time = time.time()
            output_data = model(input_data)
            end_time = time.time()
            total_time += end_time - start_time
    pytorch_fps = num_iterations / total_time
    print(f"PyTorch FPS: {pytorch_fps:.2f}")

if __name__ == "__main__":
    print("Without TensorRT")
    compute()
    print("With TensorRT")
    compute(use_tensorrt=True)

环境信息

依赖库

torch 2.0.1
torch_tensorrt 1.4.0

GPU/CUDA信息

nvcc: NVIDIA (R) Cuda compiler driver
Cuda compilation tools, release 11.5, V11.5.119
Build cuda_11.5.r11.5/compiler.30672275_0

问题排查与解决方案

1. TensorRT精度参数配置错误

你在torch.compile的options中传入的precision是torch.cuda.FloatTensor类型,但torch_tensorrt要求该参数为字符串格式(如"fp32"、"fp16"),传入Tensor类型会导致配置失效,TensorRT未真正参与优化。

修复:
将参数改为字符串格式:

options={
    "truncate_long_and_double": True,
    "precision": "fp32",  # 或"fp16",根据需求选择
    "workspace_size" : 20 << 30
},

2. 模型设备迁移顺序错误

当前代码先编译模型再移至CUDA设备,但torch_tensorrt需要模型先在CUDA上才能完成转换优化,顺序颠倒会导致编译失效。

修复:
调整代码顺序,先将模型移至CUDA再编译:

model = MyModel().to(device)  # 先迁移设备

if use_tensorrt:
    model = torch.compile(
        model,
        backend="torch_tensorrt",
        options={
            "truncate_long_and_double": True,
            "precision": "fp32",
            "workspace_size" : 20 << 30
        },
        dynamic=False,
    )

model.eval()

3. 输入尺寸导致GPU资源饱和

你的输入batch size为8192,已占满GPU显存和计算资源,此时PyTorch原生CUDA推理已达到硬件极限,TensorRT无法再提升性能。

验证/修复:
减小batch size(如改为64或128)后重新测试,小batch下TensorRT的优化效果会更明显。

4. 版本兼容性问题

torch 2.0.1与torch_tensorrt 1.4.0兼容性较差,torch_tensorrt 1.4.0主要适配PyTorch 1.13.x版本,对PyTorch 2.x的torch.compile支持不完善,可能导致TensorRT后端未被正确触发。

修复:

  • 降级PyTorch到1.13.x版本,搭配torch_tensorrt 1.4.0;
  • 升级torch_tensorrt到2.0.0及以上版本(支持PyTorch 2.x)。

5. 计时方式不准确

time.time()无法准确计时CUDA异步操作,model(input_data)返回时GPU可能还未完成计算,导致计时结果失真。

修复:
添加CUDA同步操作确保计时准确:

for i in range(num_iterations):
    start_time = time.time()
    output_data = model(input_data)
    torch.cuda.synchronize()  # 等待GPU完成计算后再计时
    end_time = time.time()
    total_time += end_time - start_time

内容的提问来源于stack exchange,提问作者M.Tailleur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 00:14:52