LibTorch 1.10 Windows与Linux CPU推理性能差异技术问询
PyTorch/LibTorch 1.10 Windows与Linux CPU推理性能差异问题
在PyTorch/LibTorch 1.10版本中,部分包含全连接层的分类器模型在Windows 10系统下的CPU推理速度显著慢于Linux系统。
本次测试的模型在PyTorch(Python 3.9)环境中构建并训练,通过Torch JIT Script导出后,使用C++版LibTorch加载调用。测试时禁用多线程,对百张虚拟图片的平均推理耗时如下:
Windows
218ms
Linux
40ms
两台设备CPU存在差异,但该性能差距无法用硬件差异解释,且Windows设备的CPU在单核心任务上理论表现更优。
复现代码片段(输入维度为256,256,1,即HWC)
模型定义Python脚本
import torch.nn as nn import torch import torch.nn.functional as F class Model(nn.Module): def __init__(self, num_classes): super().__init__() self.conv1 = nn.Conv2d(in_channels=1, out_channels=12, kernel_size=5, stride=1, padding=1) self.bn1 = nn.BatchNorm2d(12) self.conv2 = nn.Conv2d(in_channels=12, out_channels=12, kernel_size=5, stride=1, padding=1) self.bn2 = nn.BatchNorm2d(12) self.pool = nn.MaxPool2d(2,2) self.conv4 = nn.Conv2d(in_channels=12, out_channels=24, kernel_size=5, stride=1, padding=1) self.bn4 = nn.BatchNorm2d(24) self.conv5 = nn.Conv2d(in_channels=24, out_channels=24, kernel_size=5, stride=1, padding=1) self.bn5 = nn.BatchNorm2d(24) self.fc1 = nn.Linear(24*122*122, num_classes) def forward(self, input): output = F.relu(self.bn1(self.conv1(input))) output = F.relu(self.bn2(self.conv2(output))) output = self.pool(output) output = F.relu(self.bn4(self.conv4(output))) output = F.relu(self.bn5(self.conv5(output))) #print(output.shape) output = output.view(-1, 24*122*122) output = self.fc1(output) return output
模型导出Python代码
traced_script_module = torch.jit.trace(model, images) traced_script_module.save(params["model_path"] + f'CP_epoch{epoch + 1}.pt')
测速C++程序
void test() { at::set_num_threads(1); at::init_num_threads(); torch::jit::script::Module module = torch::jit::load("classifier.pt", c10::DeviceType::CPU); module.eval(); cv::Mat m = cv::Mat::ones(256, 256, CV_8UC1); torch::Tensor tensor_image = torch::from_blob(m.data, { m.rows, m.cols, m.channels() }, at::kByte); tensor_image = tensor_image.permute({ 2,0,1 }); tensor_image = tensor_image.toType(torch::kFloat32); tensor_image.to(c10::DeviceType::CPU); torch::Tensor output; auto start = std::chrono::high_resolution_clock::now(); int runs = 100; for (size_t i = 0; i < runs; i++) { output = module.forward({ tensor_image }).toTensor().detach(); } auto duration = std::chrono::duration_cast<std::chrono::milliseconds>(std::chrono::high_resolution_clock::now() - start).count(); std::cout << duration / (float) runs << std::endl; }
环境细节
- 两个平台的PyTorch均链接Intel MKL(oneAPI 2021.3.0),且均启用MKLDNN
- Windows使用编译器MSVC 14.29.30133,Linux使用gcc(SUSE Linux)7.5.0
- 默认情况下,Windows平台MKL为静态链接,Linux为动态链接
- PyTorch构建参考官方源码构建流程
内容的提问来源于stack exchange,提问作者Bastian
相关产品推荐
相关产品推荐

