使用TensorRT加速PyTorch模型推理速度未提升,求原因排查
我尝试用TensorRT提升PyTorch模型的推理速度,写了一段测试代码,但使用TensorRT和不使用的FPS都是约14.2,没有速度提升。请问问题出在哪?有没有遗漏的配置?
测试代码
import torch import argparse import time import numpy as np import torch_tensorrt # Define a simple PyTorch model class MyModel(torch.nn.Module): def __init__(self): super().__init__() self.conv1 = torch.nn.Conv2d(3, 32, kernel_size=3, stride=1, padding=1) self.relu1 = torch.nn.ReLU() self.conv2 = torch.nn.Conv2d(32, 64, kernel_size=3, stride=1, padding=1) self.relu2 = torch.nn.ReLU() self.pool = torch.nn.MaxPool2d(kernel_size=2, stride=2) self.fc1 = torch.nn.Linear(64 * 16 * 16, 512) self.relu3 = torch.nn.ReLU() self.fc2 = torch.nn.Linear(512, 10) def forward(self, x): x = self.conv1(x) x = self.relu1(x) x = self.conv2(x) x = self.relu2(x) x = self.pool(x) x = x.view(-1, 64 * 16 * 16) x = self.fc1(x) x = self.relu3(x) x = self.fc2(x) return x def compute(use_tensorrt=False): force_cpu = False useCuda = torch.cuda.is_available() and not force_cpu if useCuda: print('Using CUDA.') dtype = torch.cuda.FloatTensor ltype = torch.cuda.LongTensor device = torch.device("cuda:0") else: print('No CUDA available.') dtype = torch.FloatTensor ltype = torch.LongTensor device = torch.device("cpu") model = MyModel() input_shape = (8192, 3, 32, 32) if use_tensorrt: model = torch.compile( model, backend="torch_tensorrt", options={ "truncate_long_and_double": True, "precision": dtype, "workspace_size" : 20 << 30 }, dynamic=False, ) model = model.to(device) model.eval() num_iterations = 100 total_time = 0.0 with torch.no_grad(): input_data = torch.randn(input_shape).to(device).type(dtype) #warmup for i in range(100): output_data = model(input_data) for i in range(num_iterations): start_time = time.time() output_data = model(input_data) end_time = time.time() total_time += end_time - start_time pytorch_fps = num_iterations / total_time print(f"PyTorch FPS: {pytorch_fps:.2f}") if __name__ == "__main__": print("Without TensorRT") compute() print("With TensorRT") compute(use_tensorrt=True)
环境信息
依赖库
torch 2.0.1 torch_tensorrt 1.4.0
GPU/CUDA信息
nvcc: NVIDIA (R) Cuda compiler driver Cuda compilation tools, release 11.5, V11.5.119 Build cuda_11.5.r11.5/compiler.30672275_0
问题排查与解决方案
1. TensorRT精度参数配置错误
你在torch.compile的options中传入的precision是torch.cuda.FloatTensor类型,但torch_tensorrt要求该参数为字符串格式(如"fp32"、"fp16"),传入Tensor类型会导致配置失效,TensorRT未真正参与优化。
修复:
将参数改为字符串格式:
options={ "truncate_long_and_double": True, "precision": "fp32", # 或"fp16",根据需求选择 "workspace_size" : 20 << 30 },
2. 模型设备迁移顺序错误
当前代码先编译模型再移至CUDA设备,但torch_tensorrt需要模型先在CUDA上才能完成转换优化,顺序颠倒会导致编译失效。
修复:
调整代码顺序,先将模型移至CUDA再编译:
model = MyModel().to(device) # 先迁移设备 if use_tensorrt: model = torch.compile( model, backend="torch_tensorrt", options={ "truncate_long_and_double": True, "precision": "fp32", "workspace_size" : 20 << 30 }, dynamic=False, ) model.eval()
3. 输入尺寸导致GPU资源饱和
你的输入batch size为8192,已占满GPU显存和计算资源,此时PyTorch原生CUDA推理已达到硬件极限,TensorRT无法再提升性能。
验证/修复:
减小batch size(如改为64或128)后重新测试,小batch下TensorRT的优化效果会更明显。
4. 版本兼容性问题
torch 2.0.1与torch_tensorrt 1.4.0兼容性较差,torch_tensorrt 1.4.0主要适配PyTorch 1.13.x版本,对PyTorch 2.x的torch.compile支持不完善,可能导致TensorRT后端未被正确触发。
修复:
- 降级PyTorch到1.13.x版本,搭配torch_tensorrt 1.4.0;
- 升级torch_tensorrt到2.0.0及以上版本(支持PyTorch 2.x)。
5. 计时方式不准确
time.time()无法准确计时CUDA异步操作,model(input_data)返回时GPU可能还未完成计算,导致计时结果失真。
修复:
添加CUDA同步操作确保计时准确:
for i in range(num_iterations): start_time = time.time() output_data = model(input_data) torch.cuda.synchronize() # 等待GPU完成计算后再计时 end_time = time.time() total_time += end_time - start_time
内容的提问来源于stack exchange,提问作者M.Tailleur

