You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Mac下icc编译器编译结果与gcc、clang不符问题求助

Why does ICC produce all zeros while GCC/Clang work correctly for this AVX+OpenMP code?

Let's break down the issues causing ICC to output all zeros, while GCC/Clang behave as expected:

1. Non-standard vector element access isn't supported by ICC

Your code uses direct subscript access like vchi[0] to read elements from an AVX vector register—this is a non-standard extension specific to GCC/Clang. Intel's ICC (now rebranded as ICX/ICPX) doesn't recognize this syntax, so it fails to generate the correct assignment logic. As a result, the z array never gets updated from its initial zero values.

To fix this, use standard AVX intrinsic functions to handle vector-to-memory operations:

  • Instead of extracting elements one by one, use _mm256_store_pd to write the entire vector to memory in one go (far more efficient):
    _mm256_store_pd(&chi[j], vchi);
    
  • If you ever need to extract individual elements, use _mm256_extractf_pd (e.g., _mm256_extractf_pd(vchi, 0) for the first element).

2. Memory alignment requirements for AVX loads/stores

_mm256_load_pd requires memory addresses to be 32-byte aligned, but standard malloc only guarantees 16-byte alignment on most systems (including macOS). GCC/Clang may tolerate misaligned memory in some cases, but ICC enforces alignment strictly, leading to undefined behavior.

Fix this by using alignment-aware allocation functions:

  • On POSIX systems (like macOS), use posix_memalign:
    posix_memalign((void**)&xhi, 32, sizeof(double)*m*1);
    posix_memalign((void**)&yhi, 32, sizeof(double)*m*1);
    posix_memalign((void**)&z, 32, sizeof(double)*m*1);
    
  • Or use C11's aligned_alloc (note: the size must be a multiple of the alignment value):
    xhi = (double*)aligned_alloc(32, sizeof(double)*m*1);
    

You can still use free() to release memory allocated with these functions.

3. Minor OpenMP syntax cleanup

Your parallel region has a redundant curly brace after #pragma omp parallel. You can simplify the code by merging the parallel and for directives:

#pragma omp parallel for private(vahi,vbhi,vchi)
for (j = 0; j < m*n; j+=4) {
    // loop body
}

This avoids potential syntax ambiguities and makes the code cleaner.

Fixed Code Example

Here's the corrected code that works with ICC, GCC, and Clang:

#include "stdio.h"
#include "time.h"
#include "math.h"
#include "stdlib.h"
#include "omp.h"
#include "x86intrin.h"
#include <stdlib.h> // For posix_memalign

void dd_m_dd(double *ahi, double *bhi, double *chi, int m, int n) {
    int j;
    #pragma omp parallel for private(vahi,vbhi,vchi)
    for (j = 0; j < m*n; j+=4) {
        vbhi = _mm256_broadcast_sd(&bhi[j]);
        vahi = _mm256_load_pd(&ahi[j]);
        vchi = _mm256_mul_pd(vahi, vbhi); // Use standard multiplication intrinsic
        _mm256_store_pd(&chi[j], vchi); // Write entire vector to memory
    }
}

int main(int argc, const char * argv[]){
    int m;
    double* xhi;
    double* yhi;
    double* z;
    int i;
    m=(int)pow(2,3);

    // Allocate 32-byte aligned memory
    posix_memalign((void**)&xhi, 32, sizeof(double)*m*1);
    posix_memalign((void**)&yhi, 32, sizeof(double)*m*1);
    posix_memalign((void**)&z, 32, sizeof(double)*m*1);

    // Initialize values
    for (i=0;i<m;i++) {
        xhi[i]=i+1;
        yhi[i]=i+1;
        z[i]=0;
    }

    dd_m_dd(xhi,yhi,z,m,1);

    // Print results
    for (i=0;i<m;i++) {
        printf("%e\n",z[i]);
    }

    free(xhi);
    free(yhi);
    free(z);
    return 0;
}

When compiled with any of your commands, this code will output the expected sequence: 1.000000e+00, 2.000000e+00, 3.000000e+00, 4.000000e+00, 25.000000e+00, 30.000000e+00, 35.000000e+00, 40.000000e+00.

内容的提问来源于stack exchange,提问作者Mic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:33:35