AVX2矩阵加法中_mm256_load_ps触发段错误,_mm256_set_ps正常
问题:AVX2
_mm256_load_ps触发段错误,_mm256_set_ps可正常运行 问题背景
我正在开发一款涉及多矩阵运算的高吞吐低延迟实时程序,打算用AVX2/AVX512指令集提升性能,这是我第一次使用SIMD的AVX指令集,基于g++提供的AVX Intrinsics函数开发。当前遇到的问题:使用_mm256_load_ps会触发段错误,但_mm256_set_ps能正常运行。由于_mm256_load_ps性能更优,需要排查问题原因。
复现代码
#include <immintrin.h> #include <string.h> const std::uint64_t MAX_COUNT = 100000; int main() { float mat1[MAX_COUNT], mat2[MAX_COUNT], rslt[MAX_COUNT]; for(int i = 0; i < MAX_COUNT; i++){ mat1[i] = i; mat2[i] = 100-i; } for(int i = 0; i < MAX_COUNT; i +=8) { // 正常运行 //auto avx_a = _mm256_set_ps(mat1[i+7], mat1[i+6], mat1[i+5], mat1[i+4], mat1[i+3], mat1[i+2], mat1[i+1], mat1[i+0]); // 正常运行 //auto avx_b = _mm256_set_ps(mat2[i+7], mat2[i+6], mat2[i+5], mat2[i+4], mat2[i+3], mat2[i+2], mat2[i+1], mat2[i+0]); // 触发段错误 auto avx_a = _mm256_load_ps(&mat1[i]); // 触发段错误 auto avx_b = _mm256_load_ps(&mat2[i]); auto avx_c = _mm256_add_ps(avx_a, avx_b); float *result = (float*)&avx_c; memcpy(&rslt[i], result, 8*sizeof(float)); } return 0; }
尝试对齐数据时的编译错误
我尝试用__declspec(align(32))对齐数组,代码如下:
__declspec(align(32)) float mat1[MAX_COUNT]
但出现编译错误:
test_2.cpp: In function ‘int main()’: test_2.cpp:11:21: error: too few arguments to function ‘void* std::align(std::size_t, std::size_t, void*&, std::size_t&)’ 11 | __declspec(align(32)) float mat1[MAX_COUNT]; | ~~~~~^~~~ In file included from /usr/include/c++/11/memory:72, from /usr/include/x86_64-linux-gnu/c++/11/bits/stdc++.h:82, from test_2.cpp:2: /usr/include/c++/11/bits/align.h:62:1: note: declared here 62 | align(size_t __align, size_t __size, void*& __ptr, size_t& __space) noexcept | ^~~~~ test_2.cpp:11:5: error: ‘__declspec’ was not declared in this scope 11 | __declspec(align(32)) float mat1[MAX_COUNT]; | ^~~~~~~~~~
问题原因与解决方法
段错误根源
_mm256_load_ps要求内存地址必须是32字节对齐的(AVX2的256位向量对应32字节)。而栈上默认分配的数组(比如代码中的mat1、mat2)无法保证满足32字节对齐要求,当&mat1[i]的地址不对齐时,执行_mm256_load_ps就会触发段错误。
_mm256_set_ps是逐个加载单个float值构造向量,不要求内存对齐,所以能正常运行,但性能远低于直接加载对齐内存的_mm256_load_ps。
解决对齐问题(g++环境)
__declspec(align(32))是MSVC专属语法,g++不支持,需改用__attribute__((aligned(32)))指定内存对齐:
修改数组定义部分:
__attribute__((aligned(32))) float mat1[MAX_COUNT]; __attribute__((aligned(32))) float mat2[MAX_COUNT]; __attribute__((aligned(32))) float rslt[MAX_COUNT];
另外注意:当前MAX_COUNT=100000刚好是8的倍数,循环不会越界;如果后续修改为非8的倍数,需要用标量运算处理剩余的元素。
额外优化建议
- 不要用
memcpy复制向量结果,改用_mm256_store_ps(&rslt[i], avx_c),性能更高,且同样要求rslt[i]地址对齐(所以rslt也需要对齐)。 - 编译时必须加上AVX2编译选项:
g++ -mavx2 your_code.cpp -o your_program,确保编译器生成AVX2指令。
内容的提问来源于stack exchange,提问作者Dark Sorrow
相关产品推荐
相关产品推荐

