为何Intel i3-N305的SIMD索引最大值函数性能比AMD Ryzen7 3800X慢3.2倍?
我在Intel i3-N305(3.8GHz)和AMD Ryzen 7 3800X(3.9GHz)主机上运行了由gcc-13编译的同一二进制文件,代码使用VCL库实现,两者性能差距达3.2倍,远超预期。已知两款CPU单线程SIMD能力相近,且7-zip基准测试表现接近,求解释原因。
测试代码
int loop_vc_nested(const array<uint8_t, H*W> &img, const array<Vec32uc, 8> &idx) { int sum = 0; Vec32uc vMax, iMax, vCurr, iCurr; for (int i=0; i<H*W; i+=W) { iMax.load(&idx[0]); vMax.load(&img[i]); for (int j=1; j<8; j++) { iCurr.load(&idx[j]); vCurr.load(&img[i+j*32]); iMax = select(vCurr > vMax, iCurr, iMax); vMax = max(vMax, vCurr); } Vec32uc vMaxAll{horizontal_max(vMax)}; sum += iMax[horizontal_find_first(vMax == vMaxAll)]; } return sum; }
编译选项:-O3 -Wno-narrowing -ffast-math -fno-trapping-math -fno-math-errno -ffinite-math-only -march=alderlake
基准测试耗时
AMD Ryzen 7 3800X(Ubuntu 22.04.3 LTS,gcc 13.1)
loop_vc_nested(): 3.597 3.777 [us] 108834
Intel i3-N305(Ubuntu 23.10,gcc 13.1)
loop_vc_nested(): 11.804 11.922 [us] 108834
Perf性能分析数据
AMD Ryzen 7 3800X
3,841.61 msec task-clock # 1.000 CPUs utilized 20 context-switches # 5.206 /sec 0 cpu-migrations # 0.000 /sec 2,191 page-faults # 570.333 /sec 14,909,837,582 cycles # 3.881 GHz (83.34%) 3,509,824 stalled-cycles-frontend # 0.02% frontend cycles idle (83.34%) 9,865,497,290 stalled-cycles-backend # 66.17% backend cycles idle (83.34%) 42,856,816,868 instructions # 2.87 insn per cycle # 0.23 stalled cycles per insn (83.34%) 1,718,672,677 branches # 447.383 M/sec (83.34%) 2,409,251 branch-misses # 0.14% of all branches (83.29%)
Intel i3-N305
12,015.18 msec task-clock # 1.000 CPUs utilized 57 context-switches # 4.744 /sec 0 cpu-migrations # 0.000 /sec 2,196 page-faults # 182.769 /sec 45,432,594,158 cycles # 3.781 GHz (74.97%) 42,847,054,707 instructions # 0.94 insn per cycle (87.48%) 1,714,003,765 branches # 142.653 M/sec (87.48%) 4,254,872 branch-misses # 0.25% of all branches (87.51%) TopdownL1 # 0.2 % tma_bad_speculation # 45.5 % tma_retiring (87.52%) # 53.8 % tma_backend_bound # 53.8 % tma_backend_bound_aux # 0.5 % tma_frontend_bound (87.52%)
Intel i3-N305缓存使用信息(perf stat -d)
15,615,324,576 L1-dcache-loads # 1.294 G/sec (54.50%) <not supported> L1-dcache-load-misses 60,909 LLC-loads # 5.048 K/sec (54.50%) 5,231 LLC-load-misses # 8.59% of all L1-icache accesses (54.50%)
Intel编译器(icx)测试结果
Ubuntu 23.10 on Intel(R) Core(TM) i3-N305
clang v17.0.0 (icx 2024.0.2.20231213) C++
loop_vc_nested(): 12.311 12.397 [us] 108834 # -march=native
loop_vc_nested(): 12.773 12.847 [us] 108834 # -march=alderlake
loop_vc_nested(): 12.418 12.519 [us] 108834 # -march=gracemont
loop_vc_unrolled(): 10.388 12.406 [us] 108834 # -march=gracemont
loop_vc_nested_noselect_2chains(): 6.686 10.454 [us] 109599 # -march=gracemont
性能差距根源分析
从perf数据可以看出核心差异:
- IPC(每周期指令数)悬殊:Ryzen 3800X的IPC达2.87,而i3-N305仅0.94,说明N305的指令执行效率极低。
- 后端瓶颈严重:N305的Topdown分析显示53.8%的周期处于后端绑定状态,流水线阻塞严重;而Ryzen虽然后端闲置率高,但IPC仍远超N305。
- 内存访问效率差异:N305的L1缓存负载量巨大(1.294G/sec),结合其低功耗架构的内存带宽、缓存延迟劣势,推测L1缓存访问延迟导致SIMD指令无法高效流水执行。
结合CPU微架构特性:
- Ryzen 3800X基于Zen2,拥有2个256位SIMD执行单元,内存子系统带宽较高,能高效支撑持续的SIMD内存访问与计算。
- i3-N305基于Gracemont低功耗架构,虽支持AVX2,但SIMD执行单元宽度仅128位(256位指令需拆分为两次执行),且乱序执行窗口较小,难以掩盖内存访问延迟,进一步放大了后端瓶颈。
另外,编译选项-march=alderlake针对混合架构优化,未充分适配Gracemont的微架构特性;尝试-march=gracemont后性能无明显提升,说明gcc对该架构的SIMD代码生成优化仍有不足。
内容的提问来源于stack exchange,提问作者Paul Jurczak

