You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

不同CPU配置下Donut模型Linux环境推理耗时差异过大的原因排查

推理耗时差异排查:Donut模型在两台Linux服务器上的性能问题

我在两台Linux服务器上运行以下Donut模型代码,出现了巨大的推理耗时差异:

from donut import DonutModel
import time
from PIL import Image
import torch
model = DonutModel.from_pretrained("naver-clova-ix/donut-base-finetuned-cord-v2")

model.encoder.to(torch.bfloat16)
model.eval() 
image = Image.open("./donut/misc/sample_image_cord_test_receipt_00004.png").convert("RGB")
t=time.time()
output = model.inference(image=image, prompt="<s_cord-v2>")
print(time.time()-t)
print(output)
  • 服务器A:Intel(R) Xeon(R) CPU E5-2640 v4 @2.40GHz(40线程、62Gi内存),推理耗时815.7秒
  • 服务器B:Intel(R) Xeon(R) Silver 4216 CPU @2.10GHz(64线程、125Gi内存),推理耗时仅约6秒

调试确认卡顿出现在model.inference阶段,以下是两台服务器的详细硬件参数:

服务器A硬件参数

CPU信息

$ lscpu
Architecture:            x86_64
  CPU op-mode(s):        32-bit, 64-bit
  Address sizes:         46 bits physical, 48 bits virtual
  Byte Order:            Little Endian
CPU(s):                  40
  On-line CPU(s) list:   0-39
Vendor ID:               GenuineIntel
  Model name:            Intel(R) Xeon(R) CPU E5-2640 v4 @ 2.40GHz
    CPU family:          6
    Model:               79
    Thread(s) per core:  2
    Core(s) per socket:  10
    Socket(s):           2
    Stepping:            1
    CPU max MHz:         3400.0000
    CPU min MHz:         1200.0000
    BogoMIPS:            4800.25
    Flags:               fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi
                          mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon p
                         ebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monito
                         r ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic mov
                         be popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_f
                         ault epb cat_l3 cdp_l3 invpcid_single pti ssbd ibrs ibpb stibp tpr_shadow vnmi flexprior
                         ity ept vpid ept_ad fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm cqm rdt
                         _a rdseed adx smap intel_pt xsaveopt cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local d
                         therm ida arat pln pts md_clear flush_l1d
Virtualization features:
  Virtualization:        VT-x
Caches (sum of all):
  L1d:                   640 KiB (20 instances)
  L1i:                   640 KiB (20 instances)
  L2:                    5 MiB (20 instances)
  L3:                    50 MiB (2 instances)
NUMA:
  NUMA node(s):          2
  NUMA node0 CPU(s):     0,2,4,6,8,10,12,14,16,18,20,22,24,26,28,30,32,34,36,38
  NUMA node1 CPU(s):     1,3,5,7,9,11,13,15,17,19,21,23,25,27,29,31,33,35,37,39
Vulnerabilities:
  Itlb multihit:         KVM: Mitigation: VMX disabled
  L1tf:                  Mitigation; PTE Inversion; VMX conditional cache flushes, SMT vulnerable
  Mds:                   Mitigation; Clear CPU buffers; SMT vulnerable
  Meltdown:              Mitigation; PTI
  Mmio stale data:       Mitigation; Clear CPU buffers; SMT vulnerable
  Retbleed:              Not affected
  Spec store bypass:     Mitigation; Speculative Store Bypass disabled via prctl and seccomp
  Spectre v1:            Mitigation; usercopy/swapgs barriers and __user pointer sanitization
  Spectre v2:            Mitigation; Retpolines, IBPB conditional, IBRS_FW, STIBP conditional, RSB filling, PBRSB
                         -eIBRS Not affected
  Srbds:                 Not affected
  Tsx async abort:       Mitigation; Clear CPU buffers; SMT vulnerable

内存信息

$ free -h 
               total        used        free      shared  buff/cache   available
Mem:            62Gi       6.4Gi        42Gi       7.0Mi        14Gi        55Gi
Swap:          8.0Gi          0B       8.0Gi

服务器B硬件参数

CPU信息

$ lscpu
Architecture:           x86_64
  CPU op-mode(s):       32-bit, 64-bit
  Address sizes:        46 bits physical, 48 bits virtual
  Byte Order:           Little Endian
CPU(s):                 64
  On-line CPU(s) list:  0-63
Vendor ID:              GenuineIntel
  Model name:           Intel(R) Xeon(R) Silver 4216 CPU @ 2.10GHz
    CPU family:         6
    Model:              85
    Thread(s) per core: 2
    Core(s) per socket: 16
    Socket(s):          2
    Stepping:           7
    CPU max MHz:        3200.0000
    CPU min MHz:        800.0000
    BogoMIPS:           4200.00
    Flags:              fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxs
                        r sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_
                        good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl smx est tm2
                         ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes
                         xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 invpcid_single in
                        tel_ppin ssbd mba ibrs ibpb stibp ibrs_enhanced fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms inv
                        pcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw a
                        vx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm
                        ida arat pln pts pku ospke avx512_vnni md_clear flush_l1d arch_capabilities
Caches (sum of all):
  L1d:                  1 MiB (32 instances)
  L1i:                  1 MiB (32 instances)
  L2:                   32 MiB (32 instances)
  L3:                   44 MiB (2 instances)
NUMA:
  NUMA node(s):         2
  NUMA node0 CPU(s):    0,2,4,6,8,10,12,14,16,18,20,22,24,26,28,30,32,34,36,38,40,42,44,46,48,50,52,54,56,58,60,62
  NUMA node1 CPU(s):    1,3,5,7,9,11,13,15,17,19,21,23,25,27,29,31,33,35,37,39,41,43,45,47,49,51,53,55,57,59,61,63
Vulnerabilities:
  Itlb multihit:        KVM: Mitigation: VMX unsupported
  L1tf:                 Not affected
  Mds:                  Not affected
  Meltdown:             Not affected
  Mmio stale data:      Mitigation; Clear CPU buffers; SMT vulnerable
  Retbleed:             Mitigation; Enhanced IBRS
  Spec store bypass:    Mitigation; Speculative Store Bypass disabled via prctl and seccomp
  Spectre v1:           Mitigation; usercopy/swapgs barriers and __user pointer sanitization
  Spectre v2:           Mitigation; Enhanced IBRS, IBPB conditional, RSB filling, PBRSB-eIBRS SW sequence
  Srbds:                Not affected
  Tsx async abort:      Mitigation; TSX disabled

内存信息

$ free -h 
     total        used        free      shared  buff/cache   available
Mem:           125Gi       5.5Gi        51Gi       6.0Mi        68Gi       118Gi
Swap:          8.0Gi       116Mi       7.9Gi

可能的耗时差异原因分析

  • CPU指令集差异:服务器B的Xeon Silver 4216支持AVX-512指令集(从flags里的avx512f/avx512dq等可以看出),而服务器A的Xeon E5-2640 v4仅支持AVX2。Donut模型基于Transformer架构,大量矩阵运算能被AVX-512大幅加速,这是性能差异的核心原因之一。
  • PyTorch优化库兼容性:服务器B可能安装了针对Intel CPU优化的PyTorch版本(如启用Intel MKL、OneDNN加速),而服务器A的PyTorch未配置这些优化,导致矩阵运算效率低下。
  • CPU调度与频率策略:虽然服务器A基础主频更高,但服务器B的Turbo频率可达3.2GHz,且可能处于performance CPU governor模式,未被其他任务抢占资源;服务器A可能处于powersave模式,或被后台进程占用大量CPU资源。
  • 缓存与内存性能:服务器B的L2缓存总容量是服务器A的6.4倍(32MiB vs 5MiB),单核心L2缓存也更大(1MiB vs 256KiB),能大幅减少内存访问延迟。另外,若服务器A的NUMA配置不合理,跨节点内存访问会进一步增加耗时。
  • 软件环境版本差异:两台服务器的PyTorch、Python、依赖库(如torchvision、PIL)版本不一致,旧版本可能存在性能瓶颈或未支持新CPU特性。
  • 低精度推理支持:代码中设置了model.encoder.to(torch.bfloat16),但服务器A的CPU不支持bfloat16硬件加速,导致实际运行在float32模式,计算量翻倍;而服务器B支持bfloat16,能高效运行低精度推理。

内容的提问来源于stack exchange,提问作者Sarah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 21:04:52