You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

内联汇编实现的std::ops::Add<i8>比原生慢4倍,求排查原因

性能优化问题:封装数值类型的Rust crate与原生类型的性能差异

问题背景

我开发了一个封装数值原语的Rust crate,正在testing分支中对比该crate与原生core::primitive类型的性能差异,发现自定义类型的加法性能远低于原生类型,即使更换加法实现方式也无明显改善,想找出问题所在。

基准测试代码

use criterion::{criterion_group, criterion_main, Criterion};

use criterion::BenchmarkId;
use rand::Rng;
pub fn criterion_benchmark(c: &mut Criterion) {
    let mut group = c.benchmark_group("add");
    
    let mut rng = rand::thread_rng();

    // Reduce warmup and measurement time so the benchmarks don't take as long.
    group.warm_up_time(std::time::Duration::from_millis(500));
    group.measurement_time(std::time::Duration::from_millis(1000));

    group.bench_with_input(BenchmarkId::new("core", 8), &(), |b,_| {
        let lhs = rng.gen_range(0..i8::MAX/4);
        let rhs = rng.gen_range(0..i8::MAX/4);
        b.iter(|| lhs+rhs)
    });

    group.bench_with_input(BenchmarkId::new("ux2", 8), &(), |b,_| {
        let lhs = ux2::i7::try_from(rng.gen_range(0..i8::MAX/4)).unwrap();
        let rhs = ux2::i7::try_from(rng.gen_range(0..i8::MAX/4)).unwrap();
        b.iter(|| lhs+rhs)
    });

}

criterion_group!(benches, criterion_benchmark);
criterion_main!(benches);

基准测试结果

jonathan@jonathan-System-Product-Name:~/Projects/ux2/ux2$ cargo bench
   Compiling ux2 v0.7.0 (/home/jonathan/Projects/ux2/ux2)
    Finished bench [optimized] target(s) in 0.83s
     Running unittests src/lib.rs (/home/jonathan/Projects/ux2/target/release/deps/ux2-6bc559fda932b8ac)

running 0 tests

test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

     Running benches/benchmark.rs (/home/jonathan/Projects/ux2/target/release/deps/benchmark-d9482b576c6cf74e)
Gnuplot not found, using plotters backend
add/core/8              time:   [217.64 ps 221.12 ps 226.10 ps]
                        change: [-1.5311% -0.5221% +0.7828%] (p = 0.39 > 0.05)
                        No change in performance detected.
Found 23 outliers among 100 measurements (23.00%)
  1 (1.00%) low severe
  6 (6.00%) low mild
  5 (5.00%) high mild
  11 (11.00%) high severe
add/ux2/8               time:   [1.1005 ns 1.1020 ns 1.1037 ns]
                        change: [+0.9806% +1.2151% +1.4177%] (p = 0.00 < 0.05)
                        Change within noise threshold.

jonathan@jonathan-System-Product-Name:~/Projects/ux2/ux2$ 

自定义类型加法实现代码

自定义ux2::i7的Add trait由过程宏生成,实现如下:

impl std::ops::Add<ux2::i8> for ux2::i7 {
    type Output = ux2::i8;
    fn add(self, rhs: ux2::i7) -> Self::Output {
        let mut x = ux2::i8::from(self).0;
        let y = ux2::i8::from(rhs).0;
        unsafe {
            std::arch::asm! {
                "add {0}, {1}",
                inout(reg_byte) x,
                in(reg_byte) y,
            }
        }
        ux2::i8(x)
    }
}
// ...
pub struct i8(i8);
pub struct i7(i8);
impl From<i7> for i8 {
    fn from(x: i7) -> i8 {
        i8(i7.0)
    }
}

问题分析与解决思路

  • 结构体布局未做透明标注:你的自定义数值结构体i7和i8没有添加#[repr(transparent)],编译器无法将其视为与底层i8完全等价的类型,会保留结构体包装的额外开销。添加该标注后,编译器会将结构体的内存布局与内部值完全对齐,允许更多优化,比如直接跳过结构体包装操作。
  • 多余的类型转换与包装:当前加法实现中包含i7转i8、取出内部值、加法后再包装回i8的多步操作,即使这些操作看起来简单,编译器也可能无法完全消除。可以简化实现,直接访问内部值进行计算后再包装,减少中间步骤:
    impl std::ops::Add<ux2::i7> for ux2::i7 {
        type Output = ux2::i8;
        fn add(self, rhs: ux2::i7) -> Self::Output {
            ux2::i8(self.0 + rhs.0)
        }
    }
    
  • 内联汇编的反作用:手动编写内联汇编会阻止编译器进行常量折叠、指令调度等优化,对于简单加法完全没必要。移除内联汇编,让编译器生成最优指令即可。
  • 过程宏生成代码冗余:用cargo expand展开过程宏生成的代码,检查是否存在多余的临时变量、类型转换等冗余操作,优化宏的实现逻辑。
  • 基准测试的一致性:确保原生类型和自定义类型的测试场景完全对齐,比如原生测试是否也需要模拟类似的类型约束(如果有的话),避免测试条件不一致导致的性能偏差。

内容的提问来源于stack exchange,提问作者Jonathan Woollett-light

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 13:25:27