如何调整CLANG/LLVM编译选项禁用过度自动向量化?
禁用Clang/LLVM过度激进的向量化优化
我需要禁用Clang/LLVM的某类优化选项,因为它生成了臃肿的输出,核心问题是过度激进的向量化。
代码示例
typedef unsigned char uint8; typedef unsigned long ulong; typedef unsigned long long int uint64; typedef unsigned long long int size_t; ulong inline /*__attribute__((optnone))*/ SetMemberMask(uint8* const stream, size_t const length) { size_t pos = 0; uint64 compose{}; do { auto bit = stream[pos]; compose |= 1ULL << bit; } while (++pos <= length); return compose; } uint64 BitFieldStreamMasks_64(uint8* const __restrict p_stream, size_t const p_qry_rows, uint64* const __restrict p_comparand, uint64* const __restrict p_masks) { auto const bits = static_cast<size_t>(p_stream[0]); auto const memberCount = static_cast<size_t>(p_stream[1]); auto const bitOrderedIx = &p_stream[0]; auto const sectionSortedIx = &bitOrderedIx[bits]; uint64 stream = SetMemberMask(p_stream, memberCount); return stream; }
问题详情
问题集中在SetMemberMask()函数:
- 使用ICC 2021.6.0搭配编译选项
-xCORE-AVX2 -O3时,函数编译结果符合预期; - 使用Clang 15.0.0搭配
-march=x86-64-v3 -O3时,生成大量冗余向量化指令,代码臃肿; - 移除
-march=x86-64-v3或降低优化等级至-O1时,编译结果与ICC输出接近。
需求
我可以通过__attribute__((optnone))属性局部禁用该函数的优化,但希望全局禁用这类过度向量化优化。需要类似-fno-slp-vectorize的专用禁用选项,或是明确-march=x86-64-v3启用的全部优化项清单。
内容的提问来源于stack exchange,提问作者IamIC
相关产品推荐
相关产品推荐

