为何将所有小批量预加载到列表后PyTorch推理速度更快?
PyTorch DataLoader与BatchNorm性能差异问题
在4核CPU、Python 3.8环境下测试MNIST数据集时,发现PyTorch内置DataLoader存在特殊性能表现:直接迭代DataLoader做模型前向传播,比先将所有batch预加载到列表后再迭代的速度快得多,其中aten::batch_norm的耗时差距尤为显著。
测试代码如下:
import torch, torchvision import torch.nn as nn import torchvision.transforms as T from torch.profiler import profile, record_function, ProfilerActivity mnist_dataset = torchvision.datasets.MNIST(root=".", train=True, transform=T.ToTensor(), download=True) loader = torch.utils.data.DataLoader(dataset=mnist_dataset, batch_size=128,shuffle=False, pin_memory=False, num_workers=4) model = nn.Sequential(nn.Flatten(), nn.Linear(28*28, 256), nn.BatchNorm1d(256), nn.ReLU(), nn.Linear(256, 10)) model.train() with profile(activities=[ProfilerActivity.CPU], record_shapes=True) as prof: with record_function("model_inference"): for (images_iter, labels_iter) in loader: outputs_iter = model(images_iter) print(prof.key_averages().table(sort_by="cpu_time_total", row_limit=10)) with profile(activities=[ProfilerActivity.CPU], record_shapes=True) as prof: with record_function("model_inference"): train_list = [sample for sample in loader] for (images_iter, labels_iter) in train_list: outputs_iter = model(images_iter) print(prof.key_averages().table(sort_by="cpu_time_total", row_limit=10))
Torch性能分析器输出的核心部分:
直接迭代DataLoader的结果:
Name Self CPU % Self CPU CPU total % CPU total CPU time avg # of Calls aten::batch_norm 0.02% 644.000us 4.57% 134.217ms 286.177us 469 Self CPU time total: 2.937s
预加载到列表后迭代的结果:
Name Self CPU % Self CPU CPU total % CPU total CPU time avg # of Calls aten::batch_norm 70.48% 6.888s 70.62% 6.902s 14.717ms 469 Self CPU time total: 9.773s
明明是完全相同的模型前向操作,为什么直接迭代DataLoader时aten::batch_norm的耗时会大幅降低?按逻辑,预加载列表的版本因为额外的列表创建开销,整体速度应该更慢才对,这一现象无法理解。
内容的提问来源于stack exchange,提问作者grescha
相关产品推荐
相关产品推荐

