Mac下icc编译器编译结果与gcc、clang不符问题求助
Let's break down the issues causing ICC to output all zeros, while GCC/Clang behave as expected:
1. Non-standard vector element access isn't supported by ICC
Your code uses direct subscript access like vchi[0] to read elements from an AVX vector register—this is a non-standard extension specific to GCC/Clang. Intel's ICC (now rebranded as ICX/ICPX) doesn't recognize this syntax, so it fails to generate the correct assignment logic. As a result, the z array never gets updated from its initial zero values.
To fix this, use standard AVX intrinsic functions to handle vector-to-memory operations:
- Instead of extracting elements one by one, use
_mm256_store_pdto write the entire vector to memory in one go (far more efficient):_mm256_store_pd(&chi[j], vchi); - If you ever need to extract individual elements, use
_mm256_extractf_pd(e.g.,_mm256_extractf_pd(vchi, 0)for the first element).
2. Memory alignment requirements for AVX loads/stores
_mm256_load_pd requires memory addresses to be 32-byte aligned, but standard malloc only guarantees 16-byte alignment on most systems (including macOS). GCC/Clang may tolerate misaligned memory in some cases, but ICC enforces alignment strictly, leading to undefined behavior.
Fix this by using alignment-aware allocation functions:
- On POSIX systems (like macOS), use
posix_memalign:posix_memalign((void**)&xhi, 32, sizeof(double)*m*1); posix_memalign((void**)&yhi, 32, sizeof(double)*m*1); posix_memalign((void**)&z, 32, sizeof(double)*m*1); - Or use C11's
aligned_alloc(note: the size must be a multiple of the alignment value):xhi = (double*)aligned_alloc(32, sizeof(double)*m*1);
You can still use free() to release memory allocated with these functions.
3. Minor OpenMP syntax cleanup
Your parallel region has a redundant curly brace after #pragma omp parallel. You can simplify the code by merging the parallel and for directives:
#pragma omp parallel for private(vahi,vbhi,vchi) for (j = 0; j < m*n; j+=4) { // loop body }
This avoids potential syntax ambiguities and makes the code cleaner.
Fixed Code Example
Here's the corrected code that works with ICC, GCC, and Clang:
#include "stdio.h" #include "time.h" #include "math.h" #include "stdlib.h" #include "omp.h" #include "x86intrin.h" #include <stdlib.h> // For posix_memalign void dd_m_dd(double *ahi, double *bhi, double *chi, int m, int n) { int j; #pragma omp parallel for private(vahi,vbhi,vchi) for (j = 0; j < m*n; j+=4) { vbhi = _mm256_broadcast_sd(&bhi[j]); vahi = _mm256_load_pd(&ahi[j]); vchi = _mm256_mul_pd(vahi, vbhi); // Use standard multiplication intrinsic _mm256_store_pd(&chi[j], vchi); // Write entire vector to memory } } int main(int argc, const char * argv[]){ int m; double* xhi; double* yhi; double* z; int i; m=(int)pow(2,3); // Allocate 32-byte aligned memory posix_memalign((void**)&xhi, 32, sizeof(double)*m*1); posix_memalign((void**)&yhi, 32, sizeof(double)*m*1); posix_memalign((void**)&z, 32, sizeof(double)*m*1); // Initialize values for (i=0;i<m;i++) { xhi[i]=i+1; yhi[i]=i+1; z[i]=0; } dd_m_dd(xhi,yhi,z,m,1); // Print results for (i=0;i<m;i++) { printf("%e\n",z[i]); } free(xhi); free(yhi); free(z); return 0; }
When compiled with any of your commands, this code will output the expected sequence: 1.000000e+00, 2.000000e+00, 3.000000e+00, 4.000000e+00, 25.000000e+00, 30.000000e+00, 35.000000e+00, 40.000000e+00.
内容的提问来源于stack exchange,提问作者Mic

