You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整CLANG/LLVM编译选项禁用过度自动向量化?

禁用Clang/LLVM过度激进的向量化优化

我需要禁用Clang/LLVM的某类优化选项,因为它生成了臃肿的输出,核心问题是过度激进的向量化。

代码示例

typedef                     unsigned char               uint8;
typedef                     unsigned long               ulong;
typedef                     unsigned long long int      uint64;
typedef                     unsigned long long int      size_t;

ulong inline /*__attribute__((optnone))*/ SetMemberMask(uint8* const stream, size_t const length)
{
    size_t                  pos = 0;
    uint64                  compose{};

    do
    {
        auto                bit = stream[pos];

        compose |= 1ULL << bit;
    } while (++pos <= length);

    return compose;
}

uint64 BitFieldStreamMasks_64(uint8* const __restrict p_stream, size_t const p_qry_rows, uint64* const __restrict p_comparand, uint64* const __restrict p_masks)
{
    auto const              bits = static_cast<size_t>(p_stream[0]);
    auto const              memberCount = static_cast<size_t>(p_stream[1]);
    auto const              bitOrderedIx = &p_stream[0];
    auto const              sectionSortedIx = &bitOrderedIx[bits];
    uint64                  stream = SetMemberMask(p_stream, memberCount);

    return stream;
}

问题详情

问题集中在SetMemberMask()函数:

  • 使用ICC 2021.6.0搭配编译选项-xCORE-AVX2 -O3时,函数编译结果符合预期;
  • 使用Clang 15.0.0搭配-march=x86-64-v3 -O3时,生成大量冗余向量化指令,代码臃肿;
  • 移除-march=x86-64-v3或降低优化等级至-O1时,编译结果与ICC输出接近。

需求

我可以通过__attribute__((optnone))属性局部禁用该函数的优化,但希望全局禁用这类过度向量化优化。需要类似-fno-slp-vectorize的专用禁用选项,或是明确-march=x86-64-v3启用的全部优化项清单。

内容的提问来源于stack exchange,提问作者IamIC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 20:55:17