不同CPU配置下Donut模型Linux环境推理耗时差异过大的原因排查
推理耗时差异排查:Donut模型在两台Linux服务器上的性能问题
我在两台Linux服务器上运行以下Donut模型代码,出现了巨大的推理耗时差异:
from donut import DonutModel import time from PIL import Image import torch model = DonutModel.from_pretrained("naver-clova-ix/donut-base-finetuned-cord-v2") model.encoder.to(torch.bfloat16) model.eval() image = Image.open("./donut/misc/sample_image_cord_test_receipt_00004.png").convert("RGB") t=time.time() output = model.inference(image=image, prompt="<s_cord-v2>") print(time.time()-t) print(output)
- 服务器A:Intel(R) Xeon(R) CPU E5-2640 v4 @2.40GHz(40线程、62Gi内存),推理耗时815.7秒
- 服务器B:Intel(R) Xeon(R) Silver 4216 CPU @2.10GHz(64线程、125Gi内存),推理耗时仅约6秒
调试确认卡顿出现在model.inference阶段,以下是两台服务器的详细硬件参数:
服务器A硬件参数
CPU信息
$ lscpu Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 46 bits physical, 48 bits virtual Byte Order: Little Endian CPU(s): 40 On-line CPU(s) list: 0-39 Vendor ID: GenuineIntel Model name: Intel(R) Xeon(R) CPU E5-2640 v4 @ 2.40GHz CPU family: 6 Model: 79 Thread(s) per core: 2 Core(s) per socket: 10 Socket(s): 2 Stepping: 1 CPU max MHz: 3400.0000 CPU min MHz: 1200.0000 BogoMIPS: 4800.25 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon p ebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monito r ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic mov be popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_f ault epb cat_l3 cdp_l3 invpcid_single pti ssbd ibrs ibpb stibp tpr_shadow vnmi flexprior ity ept vpid ept_ad fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm cqm rdt _a rdseed adx smap intel_pt xsaveopt cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local d therm ida arat pln pts md_clear flush_l1d Virtualization features: Virtualization: VT-x Caches (sum of all): L1d: 640 KiB (20 instances) L1i: 640 KiB (20 instances) L2: 5 MiB (20 instances) L3: 50 MiB (2 instances) NUMA: NUMA node(s): 2 NUMA node0 CPU(s): 0,2,4,6,8,10,12,14,16,18,20,22,24,26,28,30,32,34,36,38 NUMA node1 CPU(s): 1,3,5,7,9,11,13,15,17,19,21,23,25,27,29,31,33,35,37,39 Vulnerabilities: Itlb multihit: KVM: Mitigation: VMX disabled L1tf: Mitigation; PTE Inversion; VMX conditional cache flushes, SMT vulnerable Mds: Mitigation; Clear CPU buffers; SMT vulnerable Meltdown: Mitigation; PTI Mmio stale data: Mitigation; Clear CPU buffers; SMT vulnerable Retbleed: Not affected Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl and seccomp Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Spectre v2: Mitigation; Retpolines, IBPB conditional, IBRS_FW, STIBP conditional, RSB filling, PBRSB -eIBRS Not affected Srbds: Not affected Tsx async abort: Mitigation; Clear CPU buffers; SMT vulnerable
内存信息
$ free -h total used free shared buff/cache available Mem: 62Gi 6.4Gi 42Gi 7.0Mi 14Gi 55Gi Swap: 8.0Gi 0B 8.0Gi
服务器B硬件参数
CPU信息
$ lscpu Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 46 bits physical, 48 bits virtual Byte Order: Little Endian CPU(s): 64 On-line CPU(s) list: 0-63 Vendor ID: GenuineIntel Model name: Intel(R) Xeon(R) Silver 4216 CPU @ 2.10GHz CPU family: 6 Model: 85 Thread(s) per core: 2 Core(s) per socket: 16 Socket(s): 2 Stepping: 7 CPU max MHz: 3200.0000 CPU min MHz: 800.0000 BogoMIPS: 4200.00 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxs r sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_ good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 invpcid_single in tel_ppin ssbd mba ibrs ibpb stibp ibrs_enhanced fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms inv pcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw a vx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts pku ospke avx512_vnni md_clear flush_l1d arch_capabilities Caches (sum of all): L1d: 1 MiB (32 instances) L1i: 1 MiB (32 instances) L2: 32 MiB (32 instances) L3: 44 MiB (2 instances) NUMA: NUMA node(s): 2 NUMA node0 CPU(s): 0,2,4,6,8,10,12,14,16,18,20,22,24,26,28,30,32,34,36,38,40,42,44,46,48,50,52,54,56,58,60,62 NUMA node1 CPU(s): 1,3,5,7,9,11,13,15,17,19,21,23,25,27,29,31,33,35,37,39,41,43,45,47,49,51,53,55,57,59,61,63 Vulnerabilities: Itlb multihit: KVM: Mitigation: VMX unsupported L1tf: Not affected Mds: Not affected Meltdown: Not affected Mmio stale data: Mitigation; Clear CPU buffers; SMT vulnerable Retbleed: Mitigation; Enhanced IBRS Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl and seccomp Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Spectre v2: Mitigation; Enhanced IBRS, IBPB conditional, RSB filling, PBRSB-eIBRS SW sequence Srbds: Not affected Tsx async abort: Mitigation; TSX disabled
内存信息
$ free -h total used free shared buff/cache available Mem: 125Gi 5.5Gi 51Gi 6.0Mi 68Gi 118Gi Swap: 8.0Gi 116Mi 7.9Gi
可能的耗时差异原因分析
- CPU指令集差异:服务器B的Xeon Silver 4216支持AVX-512指令集(从flags里的
avx512f/avx512dq等可以看出),而服务器A的Xeon E5-2640 v4仅支持AVX2。Donut模型基于Transformer架构,大量矩阵运算能被AVX-512大幅加速,这是性能差异的核心原因之一。 - PyTorch优化库兼容性:服务器B可能安装了针对Intel CPU优化的PyTorch版本(如启用Intel MKL、OneDNN加速),而服务器A的PyTorch未配置这些优化,导致矩阵运算效率低下。
- CPU调度与频率策略:虽然服务器A基础主频更高,但服务器B的Turbo频率可达3.2GHz,且可能处于
performanceCPU governor模式,未被其他任务抢占资源;服务器A可能处于powersave模式,或被后台进程占用大量CPU资源。 - 缓存与内存性能:服务器B的L2缓存总容量是服务器A的6.4倍(32MiB vs 5MiB),单核心L2缓存也更大(1MiB vs 256KiB),能大幅减少内存访问延迟。另外,若服务器A的NUMA配置不合理,跨节点内存访问会进一步增加耗时。
- 软件环境版本差异:两台服务器的PyTorch、Python、依赖库(如torchvision、PIL)版本不一致,旧版本可能存在性能瓶颈或未支持新CPU特性。
- 低精度推理支持:代码中设置了
model.encoder.to(torch.bfloat16),但服务器A的CPU不支持bfloat16硬件加速,导致实际运行在float32模式,计算量翻倍;而服务器B支持bfloat16,能高效运行低精度推理。
内容的提问来源于stack exchange,提问作者Sarah
相关产品推荐
相关产品推荐

