LTO优化适用程序类型及指令缩减机制技术问询
LTO对Dhrystone与CoreMark性能影响的测试与问题分析
在使用Dhrystone测试DMIPS时,发现LTO(链接时优化)对结果影响显著:开启LTO后的Dhrystone性能近乎是未开启时的4倍。以下是完整测试过程及数据:
未开启LTO的测试步骤与结果
首先获取测试源码并编译:
$ wget http://www.xanthos.se/~joachim/dhrystone-src.tar.gz $ cd dhrystone-src $ aarch64-linux-gnu-gcc -O3 -funroll-all-loops --param max-inline-insns-auto=550 -static dhry21a.c dhry21b.c timers.c -o dhrystone # 使用qemu-user执行
执行测试并采集性能数据(输入测试次数为100000000):
$ perf stat ./dhrystone
测试输出与性能统计:
Register option selected? YES
Microseconds for one run through Dhrystone: 0.2
Dhrystones per Second: 5234421.7
VAX MIPS rating = 2979.181
Performance counter stats for './dhrystone': 19,158.53 msec task-clock:u # 0.969 CPUs utilized 0 context-switches:u # 0.000 /sec 0 cpu-migrations:u # 0.000 /sec 547 page-faults:u # 28.551 /sec 81,470,643,102 cycles:u # 4.252 GHz (50.01%) 3,046,747 stalled-cycles-frontend:u # 0.00% frontend cycles idle (50.02%) 37,208,106,969 stalled-cycles-backend:u # 45.67% backend cycles idle (50.00%) 319,848,969,156 instructions:u # 3.93 insn per cycle # 0.12 stalled cycles per insn (49.99%) 49,311,879,609 branches:u # 2.574 G/sec (49.98%) 317,518 branch-misses:u # 0.00% of all branches (50.00%) 19.762244278 seconds time elapsed 19.118127000 seconds user 0.004017000 seconds sys
开启LTO的测试步骤与结果
编译时添加-flto参数开启链接时优化:
$ aarch64-linux-gnu-gcc -O3 -funroll-all-loops --param max-inline-insns-auto=550 -static dhry21a.c dhry21b.c timers.c -o dhrystone -flto
同样执行测试并采集性能数据:
$ perf stat ./dhrystone # 输入测试次数为100000000
测试输出与性能统计:
Register option selected? YES
Microseconds for one run through Dhrystone: 0.1
Dhrystones per Second: 19539623.0
VAX MIPS rating = 11121.015
Performance counter stats for './dhrystone': 5,146.69 msec task-clock:u # 0.908 CPUs utilized 0 context-switches:u # 0.000 /sec 0 cpu-migrations:u # 0.000 /sec 553 page-faults:u # 107.448 /sec 21,453,263,692 cycles:u # 4.168 GHz (50.00%) 1,574,543 stalled-cycles-frontend:u # 0.01% frontend cycles idle (50.03%) 12,575,396,819 stalled-cycles-backend:u # 58.62% backend cycles idle (50.04%) 89,186,371,586 instructions:u # 4.16 insn per cycle # 0.14 stalled cycles per insn (50.00%) 7,717,732,872 branches:u # 1.500 G/sec (49.97%) 353,303 branch-misses:u # 0.00% of all branches (49.96%) 5.666446006 seconds time elapsed 5.133037000 seconds user 0.003322000 seconds sys
测试结果对比
- 开启LTO的Dhrystone DMIPS为19539623.0,未开启时为5234421.7
- 开启LTO的Dhrystone执行89186371586条指令,未开启时执行319848969156条指令
从数据可见,LTO大幅减少了运行时指令数量,直接提升了Dhrystone的运行速度。但在测试CoreMark/CoreMark-Pro等基准程序时,LTO并未带来显著性能提升。
技术问题
- 哪些类型的程序更易受LTO优化影响?为何LTO对Dhrystone影响显著,对CoreMark/CoreMark-Pro却无明显作用?
- LTO是通过何种机制减少运行时指令数量的?
内容的提问来源于stack exchange,提问作者Li Chen
相关产品推荐
相关产品推荐

