You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LTO优化适用程序类型及指令缩减机制技术问询

LTO对Dhrystone与CoreMark性能影响的测试与问题分析

在使用Dhrystone测试DMIPS时,发现LTO(链接时优化)对结果影响显著:开启LTO后的Dhrystone性能近乎是未开启时的4倍。以下是完整测试过程及数据:

未开启LTO的测试步骤与结果

首先获取测试源码并编译:

$ wget http://www.xanthos.se/~joachim/dhrystone-src.tar.gz
$ cd dhrystone-src
$ aarch64-linux-gnu-gcc -O3 -funroll-all-loops --param max-inline-insns-auto=550 -static  dhry21a.c dhry21b.c timers.c -o dhrystone # 使用qemu-user执行

执行测试并采集性能数据(输入测试次数为100000000):

$ perf stat ./dhrystone

测试输出与性能统计:

Register option selected? YES
Microseconds for one run through Dhrystone: 0.2
Dhrystones per Second: 5234421.7
VAX MIPS rating = 2979.181

Performance counter stats for './dhrystone':

         19,158.53 msec task-clock:u              #    0.969 CPUs utilized          
                 0      context-switches:u        #    0.000 /sec                    
                 0      cpu-migrations:u          #    0.000 /sec                    
               547      page-faults:u             #   28.551 /sec                    
    81,470,643,102      cycles:u                  #    4.252 GHz                      (50.01%)
         3,046,747      stalled-cycles-frontend:u #    0.00% frontend cycles idle     (50.02%)
    37,208,106,969      stalled-cycles-backend:u  #   45.67% backend cycles idle      (50.00%)
   319,848,969,156      instructions:u            #    3.93  insn per cycle         
                                                  #    0.12  stalled cycles per insn  (49.99%)
    49,311,879,609      branches:u                #    2.574 G/sec                    (49.98%)
           317,518      branch-misses:u           #    0.00% of all branches          (50.00%)

      19.762244278 seconds time elapsed

      19.118127000 seconds user
       0.004017000 seconds sys

开启LTO的测试步骤与结果

编译时添加-flto参数开启链接时优化:

$ aarch64-linux-gnu-gcc -O3 -funroll-all-loops --param max-inline-insns-auto=550 -static  dhry21a.c dhry21b.c timers.c -o dhrystone -flto

同样执行测试并采集性能数据:

$ perf stat ./dhrystone # 输入测试次数为100000000

测试输出与性能统计:

Register option selected? YES
Microseconds for one run through Dhrystone: 0.1
Dhrystones per Second: 19539623.0
VAX MIPS rating = 11121.015

Performance counter stats for './dhrystone':

          5,146.69 msec task-clock:u              #    0.908 CPUs utilized          
                 0      context-switches:u        #    0.000 /sec                    
                 0      cpu-migrations:u          #    0.000 /sec                    
               553      page-faults:u             #  107.448 /sec                    
    21,453,263,692      cycles:u                  #    4.168 GHz                      (50.00%)
         1,574,543      stalled-cycles-frontend:u #    0.01% frontend cycles idle     (50.03%)
    12,575,396,819      stalled-cycles-backend:u  #   58.62% backend cycles idle      (50.04%)
    89,186,371,586      instructions:u            #    4.16  insn per cycle         
                                                  #    0.14  stalled cycles per insn  (50.00%)
     7,717,732,872      branches:u                #    1.500 G/sec                    (49.97%)
           353,303      branch-misses:u           #    0.00% of all branches          (49.96%)

       5.666446006 seconds time elapsed

       5.133037000 seconds user
       0.003322000 seconds sys

测试结果对比

  • 开启LTO的Dhrystone DMIPS为19539623.0,未开启时为5234421.7
  • 开启LTO的Dhrystone执行89186371586条指令,未开启时执行319848969156条指令

从数据可见,LTO大幅减少了运行时指令数量,直接提升了Dhrystone的运行速度。但在测试CoreMark/CoreMark-Pro等基准程序时,LTO并未带来显著性能提升。

技术问题

  1. 哪些类型的程序更易受LTO优化影响?为何LTO对Dhrystone影响显著,对CoreMark/CoreMark-Pro却无明显作用?
  2. LTO是通过何种机制减少运行时指令数量的?

内容的提问来源于stack exchange,提问作者Li Chen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 00:55:26