You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AVX2矩阵加法中_mm256_load_ps触发段错误,_mm256_set_ps正常

问题:AVX2 _mm256_load_ps触发段错误,_mm256_set_ps可正常运行

问题背景

我正在开发一款涉及多矩阵运算的高吞吐低延迟实时程序,打算用AVX2/AVX512指令集提升性能,这是我第一次使用SIMD的AVX指令集,基于g++提供的AVX Intrinsics函数开发。当前遇到的问题:使用_mm256_load_ps会触发段错误,但_mm256_set_ps能正常运行。由于_mm256_load_ps性能更优,需要排查问题原因。

复现代码

#include <immintrin.h>
#include <string.h>

const std::uint64_t MAX_COUNT = 100000;
int main()    
{
    float mat1[MAX_COUNT], mat2[MAX_COUNT], rslt[MAX_COUNT];
    for(int i = 0; i < MAX_COUNT; i++){
        mat1[i] = i;
        mat2[i] = 100-i;
    }
    
    for(int i = 0; i < MAX_COUNT; i +=8)
    {
        // 正常运行
        //auto avx_a = _mm256_set_ps(mat1[i+7], mat1[i+6], mat1[i+5], mat1[i+4], mat1[i+3], mat1[i+2], mat1[i+1], mat1[i+0]);
        // 正常运行
        //auto avx_b = _mm256_set_ps(mat2[i+7], mat2[i+6], mat2[i+5], mat2[i+4], mat2[i+3], mat2[i+2], mat2[i+1], mat2[i+0]);
        // 触发段错误
        auto avx_a = _mm256_load_ps(&mat1[i]);
        // 触发段错误
        auto avx_b = _mm256_load_ps(&mat2[i]);
        auto avx_c = _mm256_add_ps(avx_a, avx_b);
        float *result = (float*)&avx_c;
        memcpy(&rslt[i], result, 8*sizeof(float));
    }
    
    return 0;
}

尝试对齐数据时的编译错误

我尝试用__declspec(align(32))对齐数组,代码如下:

__declspec(align(32)) float mat1[MAX_COUNT]

但出现编译错误:

test_2.cpp: In function ‘int main()’:
test_2.cpp:11:21: error: too few arguments to function ‘void* std::align(std::size_t, std::size_t, void*&, std::size_t&)’
   11 |     __declspec(align(32)) float mat1[MAX_COUNT];
      |                ~~~~~^~~~
In file included from /usr/include/c++/11/memory:72,
                 from /usr/include/x86_64-linux-gnu/c++/11/bits/stdc++.h:82,
                 from test_2.cpp:2:
/usr/include/c++/11/bits/align.h:62:1: note: declared here
   62 | align(size_t __align, size_t __size, void*& __ptr, size_t& __space) noexcept
      | ^~~~~
test_2.cpp:11:5: error: ‘__declspec’ was not declared in this scope
   11 |     __declspec(align(32)) float mat1[MAX_COUNT];
      |     ^~~~~~~~~~

问题原因与解决方法

段错误根源

_mm256_load_ps要求内存地址必须是32字节对齐的(AVX2的256位向量对应32字节)。而栈上默认分配的数组(比如代码中的mat1、mat2)无法保证满足32字节对齐要求,当&mat1[i]的地址不对齐时,执行_mm256_load_ps就会触发段错误。

_mm256_set_ps是逐个加载单个float值构造向量,不要求内存对齐,所以能正常运行,但性能远低于直接加载对齐内存的_mm256_load_ps。

解决对齐问题(g++环境)

__declspec(align(32))是MSVC专属语法,g++不支持,需改用__attribute__((aligned(32)))指定内存对齐:

修改数组定义部分:

__attribute__((aligned(32))) float mat1[MAX_COUNT];
__attribute__((aligned(32))) float mat2[MAX_COUNT];
__attribute__((aligned(32))) float rslt[MAX_COUNT];

另外注意:当前MAX_COUNT=100000刚好是8的倍数,循环不会越界;如果后续修改为非8的倍数,需要用标量运算处理剩余的元素。

额外优化建议

  • 不要用memcpy复制向量结果,改用_mm256_store_ps(&rslt[i], avx_c),性能更高,且同样要求rslt[i]地址对齐(所以rslt也需要对齐)。
  • 编译时必须加上AVX2编译选项:g++ -mavx2 your_code.cpp -o your_program,确保编译器生成AVX2指令。

内容的提问来源于stack exchange,提问作者Dark Sorrow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 04:10:32