You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Intel i3-N305的SIMD索引最大值函数性能比AMD Ryzen7 3800X慢3.2倍?

为何Intel i3-N305与AMD Ryzen 7 3800X的SIMD代码性能差距达3.2倍?

我在Intel i3-N305(3.8GHz)和AMD Ryzen 7 3800X(3.9GHz)主机上运行了由gcc-13编译的同一二进制文件,代码使用VCL库实现,两者性能差距达3.2倍,远超预期。已知两款CPU单线程SIMD能力相近,且7-zip基准测试表现接近,求解释原因。

测试代码

int loop_vc_nested(const array<uint8_t, H*W> &img, const array<Vec32uc, 8> &idx) {
  int sum = 0;
  Vec32uc vMax, iMax, vCurr, iCurr;

  for (int i=0; i<H*W; i+=W) {
    iMax.load(&idx[0]);
    vMax.load(&img[i]);

    for (int j=1; j<8; j++) {
      iCurr.load(&idx[j]);
      vCurr.load(&img[i+j*32]);
      iMax = select(vCurr > vMax, iCurr, iMax);
      vMax = max(vMax, vCurr);
    }

    Vec32uc vMaxAll{horizontal_max(vMax)};
    sum += iMax[horizontal_find_first(vMax == vMaxAll)];
  }

  return sum;
}

编译选项:-O3 -Wno-narrowing -ffast-math -fno-trapping-math -fno-math-errno -ffinite-math-only -march=alderlake

基准测试耗时

AMD Ryzen 7 3800X(Ubuntu 22.04.3 LTS,gcc 13.1)
loop_vc_nested(): 3.597 3.777 [us] 108834

Intel i3-N305(Ubuntu 23.10,gcc 13.1)
loop_vc_nested(): 11.804 11.922 [us] 108834

Perf性能分析数据

AMD Ryzen 7 3800X

3,841.61 msec task-clock                       #    1.000 CPUs utilized              
                20      context-switches                 #    5.206 /sec                       
                 0      cpu-migrations                   #    0.000 /sec                       
             2,191      page-faults                      #  570.333 /sec                       
    14,909,837,582      cycles                           #    3.881 GHz                         (83.34%)
         3,509,824      stalled-cycles-frontend          #    0.02% frontend cycles idle        (83.34%)
     9,865,497,290      stalled-cycles-backend           #   66.17% backend cycles idle         (83.34%)
    42,856,816,868      instructions                     #    2.87  insn per cycle            
                                                  #    0.23  stalled cycles per insn     (83.34%)
     1,718,672,677      branches                         #  447.383 M/sec                       (83.34%)
         2,409,251      branch-misses                    #    0.14% of all branches             (83.29%)

Intel i3-N305

12,015.18 msec task-clock                       #    1.000 CPUs utilized              
                57      context-switches                 #    4.744 /sec                       
                 0      cpu-migrations                   #    0.000 /sec                       
             2,196      page-faults                      #  182.769 /sec                       
    45,432,594,158      cycles                           #    3.781 GHz                         (74.97%)
    42,847,054,707      instructions                     #    0.94  insn per cycle              (87.48%)
     1,714,003,765      branches                         #  142.653 M/sec                       (87.48%)
         4,254,872      branch-misses                    #    0.25% of all branches             (87.51%)
                        TopdownL1                 #      0.2 %  tma_bad_speculation    
                                                  #     45.5 %  tma_retiring             (87.52%)
                                                  #     53.8 %  tma_backend_bound      
                                                  #     53.8 %  tma_backend_bound_aux  
                                                  #      0.5 %  tma_frontend_bound       (87.52%)

Intel i3-N305缓存使用信息(perf stat -d)

15,615,324,576      L1-dcache-loads                  #    1.294 G/sec                       (54.50%)
   <not supported>      L1-dcache-load-misses                                                 
            60,909      LLC-loads                        #    5.048 K/sec                       (54.50%)
             5,231      LLC-load-misses                  #    8.59% of all L1-icache accesses   (54.50%)

Intel编译器(icx)测试结果

Ubuntu 23.10 on Intel(R) Core(TM) i3-N305
clang v17.0.0 (icx 2024.0.2.20231213) C++
loop_vc_nested(): 12.311 12.397 [us] 108834 # -march=native
loop_vc_nested(): 12.773 12.847 [us] 108834 # -march=alderlake
loop_vc_nested(): 12.418 12.519 [us] 108834 # -march=gracemont
loop_vc_unrolled(): 10.388 12.406 [us] 108834 # -march=gracemont
loop_vc_nested_noselect_2chains(): 6.686 10.454 [us] 109599 # -march=gracemont

性能差距根源分析

从perf数据可以看出核心差异:

  1. IPC(每周期指令数)悬殊:Ryzen 3800X的IPC达2.87,而i3-N305仅0.94,说明N305的指令执行效率极低。
  2. 后端瓶颈严重:N305的Topdown分析显示53.8%的周期处于后端绑定状态,流水线阻塞严重;而Ryzen虽然后端闲置率高,但IPC仍远超N305。
  3. 内存访问效率差异:N305的L1缓存负载量巨大(1.294G/sec),结合其低功耗架构的内存带宽、缓存延迟劣势,推测L1缓存访问延迟导致SIMD指令无法高效流水执行。

结合CPU微架构特性:

  • Ryzen 3800X基于Zen2,拥有2个256位SIMD执行单元,内存子系统带宽较高,能高效支撑持续的SIMD内存访问与计算。
  • i3-N305基于Gracemont低功耗架构,虽支持AVX2,但SIMD执行单元宽度仅128位(256位指令需拆分为两次执行),且乱序执行窗口较小,难以掩盖内存访问延迟,进一步放大了后端瓶颈。

另外,编译选项-march=alderlake针对混合架构优化,未充分适配Gracemont的微架构特性;尝试-march=gracemont后性能无明显提升,说明gcc对该架构的SIMD代码生成优化仍有不足。

内容的提问来源于stack exchange,提问作者Paul Jurczak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 15:45:20