You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何将所有小批量预加载到列表后PyTorch推理速度更快?

PyTorch DataLoader与BatchNorm性能差异问题

在4核CPU、Python 3.8环境下测试MNIST数据集时,发现PyTorch内置DataLoader存在特殊性能表现:直接迭代DataLoader做模型前向传播,比先将所有batch预加载到列表后再迭代的速度快得多,其中aten::batch_norm的耗时差距尤为显著。

测试代码如下:

import torch, torchvision
import torch.nn as nn
import torchvision.transforms as T
from torch.profiler import profile, record_function, ProfilerActivity


mnist_dataset = torchvision.datasets.MNIST(root=".", train=True, transform=T.ToTensor(), download=True)
loader = torch.utils.data.DataLoader(dataset=mnist_dataset, batch_size=128,shuffle=False, pin_memory=False, num_workers=4)
model = nn.Sequential(nn.Flatten(), nn.Linear(28*28, 256), nn.BatchNorm1d(256), nn.ReLU(), nn.Linear(256, 10))
model.train()


with profile(activities=[ProfilerActivity.CPU], record_shapes=True) as prof:
    with record_function("model_inference"):
        for (images_iter, labels_iter) in loader:
            outputs_iter = model(images_iter)
print(prof.key_averages().table(sort_by="cpu_time_total", row_limit=10))


with profile(activities=[ProfilerActivity.CPU], record_shapes=True) as prof:
    with record_function("model_inference"):
        train_list = [sample for sample in loader]
        for (images_iter, labels_iter) in train_list:
            outputs_iter = model(images_iter)
print(prof.key_averages().table(sort_by="cpu_time_total", row_limit=10))

Torch性能分析器输出的核心部分:
直接迭代DataLoader的结果:

Name                Self CPU %      Self CPU   CPU total %     CPU total  CPU time avg    # of Calls
aten::batch_norm         0.02%     644.000us         4.57%     134.217ms     286.177us           469
Self CPU time total: 2.937s

预加载到列表后迭代的结果:

Name                 Self CPU %      Self CPU   CPU total %     CPU total  CPU time avg    # of Calls
aten::batch_norm        70.48%        6.888s        70.62%        6.902s      14.717ms    469
Self CPU time total: 9.773s

明明是完全相同的模型前向操作,为什么直接迭代DataLoader时aten::batch_norm的耗时会大幅降低?按逻辑,预加载列表的版本因为额外的列表创建开销,整体速度应该更慢才对,这一现象无法理解。

内容的提问来源于stack exchange,提问作者grescha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 05:16:37