如何在Godbolt中禁用Rust的LLVM优化用于汇编学习
我在学习Rust与汇编相关知识,使用Godbolt编译器探索平台开展实践练习。
我编写的测试代码如下:
pub fn test() -> i32 { let a = 1; let b = 2; let c = 3; a + b + c }
我预期编译后得到包含栈空间分配、变量入栈、加法运算等完整步骤的汇编输出,类似如下结果:
example::test: subq $16, %rsp movl $1, (%rsp) movl $2, 4(%rsp) movl $3, 8(%rsp) movl (%rsp), %eax addl 4(%rsp), %eax addl 8(%rsp), %eax addq $16, %rsp retq
但实际编译得到的汇编直接返回常量计算结果,完全省略了中间操作步骤:
example::test: mov eax, 6 ret
该输出无法用于演示栈分配、加法运算等汇编逻辑,无法满足学习需求。
我当前配置的编译参数为:-Z mir-opt-level=0 -C opt-level=0 -C overflow-checks=off,查看MIR输出可见加法逻辑并未在MIR阶段被优化,MIR中完整保留了变量定义、加法操作的步骤:
// WARNING: This output format is intended for human consumers only // and is subject to change without notice. Knock yourself out. fn test() -> i32 { let mut _0: i32; // return place in scope 0 at /app/example.rs:2:18: 2:21 let _1: i32; // in scope 0 at /app/example.rs:3:9: 3:10 let mut _4: i32; // in scope 0 at /app/example.rs:6:5: 6:10 let mut _5: i32; // in scope 0 at /app/example.rs:6:5: 6:6 let mut _6: i32; // in scope 0 at /app/example.rs:6:9: 6:10 let mut _7: i32; // in scope 0 at /app/example.rs:6:13: 6:14 scope 1 { debug a => _1; // in scope 1 at /app/example.rs:3:9: 3:10 let _2: i32; // in scope 1 at /app/example.rs:4:9: 4:10 scope 2 { debug b => _2; // in scope 2 at /app/example.rs:4:9: 4:10 let _3: i32; // in scope 2 at /app/example.rs:5:9: 5:10 scope 3 { debug c => _3; // in scope 3 at /app/example.rs:5:9: 5:10 } } } bb0: { StorageLive(_1); // scope 0 at /app/example.rs:3:9: 3:10 _1 = const 1_i32; // scope 0 at /app/example.rs:3:13: 3:14 StorageLive(_2); // scope 1 at /app/example.rs:4:9: 4:10 _2 = const 2_i32; // scope 1 at /app/example.rs:4:13: 4:14 StorageLive(_3); // scope 2 at /app/example.rs:5:9: 5:10 _3 = const 3_i32; // scope 2 at /app/example.rs:5:13: 5:14 StorageLive(_4); // scope 3 at /app/example.rs:6:5: 6:10 StorageLive(_5); // scope 3 at /app/example.rs:6:5: 6:6 _5 = _1; // scope 3 at /app/example.rs:6:5: 6:6 StorageLive(_6); // scope 3 at /app/example.rs:6:9: 6:10 _6 = _2; // scope 3 at /app/example.rs:6:9: 6:10 _4 = Add(move _5, move _6); // scope 3 at /app/example.rs:6:5: 6:10 StorageDead(_6); // scope 3 at /app/example.rs:6:9: 6:10 StorageDead(_5); // scope 3 at /app/example.rs:6:9: 6:10 StorageLive(_7); // scope 3 at /app/example.rs:6:13: 6:14 _7 = _3; // scope 3 at /app/example.rs:6:13: 6:14 _0 = Add(move _4, move _7); // scope 3 at /app/example.rs:6:5: 6:14 StorageDead(_7); // scope 3 at /app/example.rs:6:13: 6:14 StorageDead(_4); // scope 3 at /app/example.rs:6:13: 6:14 StorageDead(_3); // scope 2 at /app/example.rs:7:1: 7:2 StorageDead(_2); // scope 1 at /app/example.rs:7:1: 7:2 StorageDead(_1); // scope 0 at /app/example.rs:7:1: 7:2 return; // scope 0 at /app/example.rs:7:2: 7:2 } }
但生成的LLVM IR直接返回常量6,说明加法操作在MIR转LLVM阶段被优化消除:
define i32 @_ZN7example4test17h2e9277ab15e59fbdE() unnamed_addr #0 !dbg !5 { start: ret i32 6, !dbg !10 } attributes #0 = { nonlazybind uwtable "probe-stack"="__rust_probestack" "target-cpu"="x86-64" }
我测试发现将常量存入元组时该优化不会触发,例如如下代码:
pub fn test() -> i32 { let a = (1,2,3); a.0 + a.1 + a.2 }
该代码可编译得到我预期的完整汇编输出。现咨询如何配置编译参数,彻底禁用该阶段的LLVM常量折叠优化,得到未被优化的汇编输出用于汇编学习。
可以从编译参数配置、代码写法调整两个方向解决,根据使用场景选择即可:
方案1:通过编译参数彻底关闭LLVM常量折叠
你遇到的常量折叠属于LLVM代码生成阶段的默认基础优化,即使opt-level=0也会默认开启部分简单常量折叠逻辑。在现有参数基础上追加两个参数即可彻底关闭这类优化:
-C no-prepopulate-passes:禁止LLVM前端自动填充默认优化Pass-C llvm-args=-disable-llvm-optzns:禁用LLVM所有默认优化逻辑
追加后完整编译参数如下:
-Z mir-opt-level=0 -C opt-level=0 -C overflow-checks=off -C no-prepopulate-passes -C llvm-args=-disable-llvm-optzns
使用这套参数编译后,LLVM不会再对MIR生成的原始IR做常量传播、常量折叠处理,输出的汇编会完整保留栈空间分配、变量入栈、逐次加法运算的步骤,完全符合学习需求。
提示:这套参数会关闭LLVM层所有默认插入的基础优化,输出汇编的冗余度会比较高,更适合观察编译原始生成逻辑使用。
方案2:调整测试代码绕开常量识别(无需修改编译参数)
如果不想调整过多LLVM相关参数,只需要修改测试代码,让编译器无法在编译期确定变量的常量值,就不会触发常量折叠,两种简单有效的写法:
- 使用
std::hint::black_box包裹字面量值,强制编译器将其视为运行时才确定的值:
pub fn test() -> i32 { let a = std::hint::black_box(1); let b = std::hint::black_box(2); let c = std::hint::black_box(3); a + b + c }
- 将运算值改为函数入参,单函数编译场景下编译器不会对入参做编译期常量求值:
pub fn test(a: i32, b: i32, c: i32) -> i32 { a + b + c }
这两种写法在原本的-C opt-level=0参数配置下,就能生成包含完整栈操作、运算流程的汇编输出,配置成本更低。
内容的提问来源于stack exchange,提问作者Blue7

