编译Rust代码时是否遗漏AVX512目标特性?性能异常排查
AVX512性能在显式指定target-feature后暴跌的问题
问题描述
我基于Rust实现了用AVX2和AVX512指令加速的图像合成函数,运行在AMD 7950x CPU上:
- 使用
RUSTFLAGS="-C target-cpu=native" cargo bench测试时,性能表现正常:
test overlay_using_avx2 ... bench: 483,596 ns/iter (+/- 10,006)
test overlay_using_avx512 ... bench: 317,818 ns/iter (+/- 729)
- 为了实现跨机器编译运行,我显式指定代码依赖的特性并做运行时特性检查,执行
RUSTFLAGS="-C target-feature=+avx2,+avx,+sse2,+avx512f,+avx512bw" cargo bench后,AVX512性能大幅下降:
test overlay_using_avx2 ... bench: 490,664 ns/iter (+/- 13,172)
test overlay_using_avx512 ... bench: 1,519,720 ns/iter (+/- 38,608)
我的疑问:
- 是否需要从
rustc --print target-features列表中启用其他特性? - 如何查看
target-cpu=native所启用的全部特性?
附基准测试代码(基于Nightly环境):
#![feature(stdsimd)] #![feature(test)] use std::arch::x86_64::*; unsafe fn overlay_chunk_avx2(this_chunk: &mut [u8], image_chunk: &[u8], c1: __m256i, c2: __m256i) { let this_ptr = this_chunk.as_mut_ptr() as *mut __m128i; let image_ptr = image_chunk.as_ptr() as *const __m128i; let this_argb = _mm_loadu_si128(this_ptr); let image_argb = _mm_loadu_si128(image_ptr); let this_u16 = _mm256_cvtepu8_epi16(this_argb); let image_u16 = _mm256_cvtepu8_epi16(image_argb); let image_alpha = _mm256_shuffle_epi8(image_u16, c1); let image_inv_alpha = _mm256_sub_epi8(c2, image_alpha); let this_blended = _mm256_mullo_epi16(this_u16, image_inv_alpha); let image_blended = _mm256_mullo_epi16(image_u16, image_alpha); let blended = _mm256_add_epi16(this_blended, image_blended); let divided = _mm256_srli_epi16(blended, 8); let lo_lane = _mm256_castsi256_si128(divided); let hi_lane = _mm256_extracti128_si256(divided, 1); let divided_u8 = _mm_packus_epi16(lo_lane, hi_lane); _mm_storeu_si128(this_ptr, divided_u8); } unsafe fn overlay_chunk_avx512(this_chunk: &mut [u8], image_chunk: &[u8], c1: __m512i, c2: __m512i) { let this_ptr = this_chunk.as_mut_ptr() as *mut i8; let image_ptr = image_chunk.as_ptr() as *const i8; let this_argb = _mm256_loadu_epi8(this_ptr); let image_argb = _mm256_loadu_epi8(image_ptr); let this_u16 = _mm512_cvtepu8_epi16(this_argb); let image_u16 = _mm512_cvtepu8_epi16(image_argb); let image_alpha = _mm512_shuffle_epi8(image_u16, c1); let image_inv_alpha = _mm512_sub_epi8(c2, image_alpha); let this_blended = _mm512_mullo_epi16(this_u16, image_inv_alpha); let image_blended = _mm512_mullo_epi16(image_u16, image_alpha); let blended = _mm512_add_epi16(this_blended, image_blended); let divided = _mm512_srli_epi16(blended, 8); let divided_u8 = _mm512_cvtepi16_epi8(divided); _mm256_storeu_epi8(this_ptr, divided_u8); } extern crate test; #[bench] fn overlay_using_avx2(bencher: &mut test::Bencher) { let mut frame = vec![0; 1920 * 1080 * 4]; let image = vec![0; 1920 * 1080 * 4]; let constant1 = unsafe { _mm256_set_epi8(-1, 24, -1, 24, -1, 24, -1, -1, -1, 16, -1, 16, -1, 16, -1, -1, -1, 8, -1, 8, -1, 8, -1, -1, -1, 0, -1, 0, -1, 0, -1, -1) }; let constant2 = unsafe { _mm256_set_epi8(0, -1, 0, -1, 0, -1, 1, 0, 0, -1, 0, -1, 0, -1, 1, 0, 0, -1, 0, -1, 0, -1, 1, 0, 0, -1, 0, -1, 0, -1, 1, 0) }; bencher.iter(|| { let frame_chunks = frame.chunks_exact_mut(128 / 8); let image_chunks = image.chunks_exact(128 / 8); for (frame_chunk, image_chunk) in frame_chunks.zip(image_chunks) { unsafe { overlay_chunk_avx2(frame_chunk, image_chunk, constant1, constant2); } } }); } #[bench] fn overlay_using_avx512(bencher: &mut test::Bencher) { let mut frame = vec![0; 1920 * 1080 * 4]; let image = vec![0; 1920 * 1080 * 4]; let constant1 = unsafe { _mm512_set_epi8(-1, 56, -1, 56, -1, 56, -1, -1, -1, 48, -1, 48, -1, 48, -1, -1, -1, 40, -1, 40, -1, 40, -1, -1, -1, 32, -1, 32, -1, 32, -1, -1, -1, 24, -1, 24, -1, 24, -1, -1, -1, 16, -1, 16, -1, 16, -1, -1, -1, 8, -1, 8, -1, 8, -1, -1, -1, 0, -1, 0, -1, 0, -1, -1) }; let constant2 = unsafe { _mm512_set_epi8(0, -1, 0, -1, 0, -1, 1, 0, 0, -1, 0, -1, 0, -1, 1, 0, 0, -1, 0, -1, 0, -1, 1, 0, 0, -1, 0, -1, 0, -1, 1, 0, 0, -1, 0, -1, 0, -1, 1, 0, 0, -1, 0, -1, 0, -1, 1, 0, 0, -1, 0, -1, 0, -1, 1, 0, 0, -1, 0, -1, 0, -1, 1, 0) }; bencher.iter(|| { let frame_chunks = frame.chunks_exact_mut(256 / 8); let image_chunks = image.chunks_exact(256 / 8); for (frame_chunk, image_chunk) in frame_chunks.zip(image_chunks) { unsafe { overlay_chunk_avx512(frame_chunk, image_chunk, constant1, constant2); } } }); }
解决方案
1. 查看target-cpu=native启用的特性
直接运行以下命令即可获取当前CPU下native对应的所有目标特性(根据操作系统替换对应target):
# Linux rustc --print target-features --target x86_64-unknown-linux-gnu -C target-cpu=native # Windows rustc --print target-features --target x86_64-pc-windows-msvc -C target-cpu=native # macOS rustc --print target-features --target x86_64-apple-darwin -C target-cpu=native
2. 补充缺失的关键特性
你当前仅指定了基础的AVX512特性,AMD 7950x(Zen4架构)的native模式还启用了多个对性能影响极大的AVX512扩展特性,需要补充:
avx512vl:支持AVX512指令操作256/128位向量,你的AVX512函数中用到的_mm256_loadu_epi8/_mm256_storeu_epi8依赖该特性实现高效优化avx512dq:支持双字/四字操作,为整数运算类AVX512指令提供性能加速avx512vnni:Zen4专属的向量神经网络指令,能大幅提升_mm512_mullo_epi16这类整数乘法操作的效率
修改后的RUSTFLAGS命令:
RUSTFLAGS="-C target-feature=+avx,+avx2,+sse2,+avx512f,+avx512bw,+avx512vl,+avx512dq,+avx512vnni" cargo bench
3. 额外优化建议
- 尽可能让内存访问对齐:你的代码中使用了未对齐加载
_mm256_loadu_epi8,若能将chunk地址对齐到32字节,改用_mm256_load_epi8可进一步降低内存访问开销 - 确认bench模式的优化等级:在Cargo.toml中显式指定bench模式使用O3优化(默认已开启,但可确保配置正确):
[profile.bench] opt-level = 3
内容的提问来源于stack exchange,提问作者Chris
相关产品推荐
相关产品推荐

